Stars
MPIE-Bench: Benchmark for multi-person character-consistent image editing under contact, with a six-axis evaluation protocol.
A 125-task benchmark for knowledge-mediated implicit associations and retrieval blind spots in long-term agent memory.
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.8, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM5.3, Gemma4, Llava, P…
A Simple and Universal Swarm Intelligence Engine, Predicting Anything. 简洁通用的群体智能引擎,预测万物
AI Agent prompts for deep paper reading — auto-routes to benchmark, methodology, or survey/opinion analysis frameworks. Works with Cursor, Claude Code, Codex, and OpenCode.
Open-source AI agent desktop app for Windows & macOS. One-click install Claude Code, MCP tools, and Skills — with sandbox isolation, multi-model support, and Feishu/Slack integration.
😎 Awesome list of Retrieval-Augmented Generation (RAG) applications in Generative AI.
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
[ACL 2026] WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
A-RAG: Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces. State-of-the-art RAG framework with keyword, semantic, and chunk read tools for multi-hop QA.
Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles
The AI that really does things. Any OS. Any Platform. The lobster way. 🦞
The paper list of "Memory in the Age of AI Agents: A Survey"
DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research systems and human experts. It does so by decomposing expert-wri…
A benchmark for LLMs on complicated tasks in the terminal
bloom - evaluate any behavior immediately 🌸🌱
一站式原生AI PPT生成应用,几分钟内生成一套幻灯片; 支持上传任意模板图片,上传任意素材&智能解析,一句话/大纲/页面描述自动生成PPT,口头修改指定区域、一键导出可编辑ppt、视频等 - An AI-native slides generator based on nano banana pro🍌
[ICLR 2026] LinearRAG: Linear Graph Retrieval Augmented Generation on Large-scale Corpora
ThinkDepth.ai Deep Research
Awesome curated collection of images and prompts built with the soon-to-launch Nano Banana Pro model. Browse limited early-access test cases that highlight Nano Banana Pro’s strengths in consistent…
Scaling Deep Research via Reinforcement Learning in Real-world Environments.
Salesforce Enterprise Deep Research
The official repo of GraphRAG-Bench for evaluating GraphRAG models. "When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation". (ICLR'26)
Tongyi Deep Research, the Leading Open-source Deep Research Agent

