Updated biweekly.
-
AI agents are autonomous software entities that leverage intelligence technologies such as large language models and reinforcement learning to interact with their environment and pursue defined goals.
-
They acquire and generate information through external APIs, sensors, and code execution capabilities, enabling context-aware decision-making and action selection.
-
Through self-learning and feedback loops, they continuously improve their performance, minimizing human intervention while handling complex tasks.
-
By supporting multimodal inputs and coordinating with other agents, they realize diverse cognitive functions such as dialogue, reasoning, and strategic planning.
- βοΈ "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents" [paper]
- π "From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents" [paper]
- π "Deep Research Agents: A Systematic Examination And Roadmap" [paper]
- π "Towards AI Search Paradigm" [paper]
- "Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge" [paper]
- "MMSearch-R1: Incentivizing LMMs to Search" [paper]
- "Towards Robust Fact-Checking: A Multi-Agent System with Advanced Evidence Retrieval" [paper]
- [Jun 2025] "AUTOMIND: Adaptive Knowledgeable Agent for Automated Data Science" [paper]
- π [Jun 2025] "Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents" [paper]
- [Jun 2025] "SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation" [paper]
- [Jun 2025] "SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications" [paper]
- [Jun 2025] "Towards Community-Driven Agents for Machine Learning Engineering" [paper]
- [Jun 2025] "MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement" [paper]
- "Oversight Structures for Agentic AI in Public-Sector Organizations" [paper]
- βοΈ "AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance" [paper]
- π "Application-Driven Value Alignment in Agentic AI Systems: Survey and Perspectives" [paper]
- "Intelligent Design 4.0: Paradigm Evolution Toward the Agentic AI Era" [paper]
- "Improved LLM Agents for Financial Document Question Answering" [paper]
- βοΈ "ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering" [paper]
- "Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine" [paper]
- "SV-LLM: An Agentic Approach for SoC Security Verification using Large Language Models" [paper]
- βοΈ "SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents" [paper]
- "Intelligent Design 4.0: Paradigm Evolution Toward the Agentic AI Era" [paper]
- "Managing Complex Failure Analysis Workflows with LLM-based Reasoning and Acting Agents" [paper]
- "AgenticControl: An Automated Control Design Framework Using Large Language Models" [paper]
- π "A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools" [paper]
- π "A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law" [paper]
- π "Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models" [paper]
- "Table-R1: Inference-Time Scaling for Table Reasoning" [paper]
- "Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning" [paper]
- "Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning" [paper]
- "Agent RL Scaling Law: Spontaneous Code Execution for Mathematical Problem Solving" [paper]
- "Reinforced Internal-External Knowledge Synergistic Reasoning for Efficient Adaptive Search Agent" [paper]
- "An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents" [paper]
- "Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning" [paper]
- "MIRROR: Multi-agent Intra- and Inter-Reflection for Optimized Reasoning in Tool Learning" [paper]
- "EvolveSearch: An Iterative Self-Evolving Search Agent" [paper]
- "VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection" [paper]
- "Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning" [paper]
- "RM-R1: Reward Modeling as Reasoning" [paper]
- "Reward Reasoning Model" [paper]
- "R3: Robust Rubric-Agnostic Reward Models" [paper]
- "AutoLibra: Agent Metric Induction from Open-Ended Feedback" [paper]
- "MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models (Short Version)" [paper]
- "MemEngine: A Unified and Modular Library for Developing Advanced Memory of LLM-based Agents" [paper]
- "MARK: Memory Augmented Refinement of Knowledge" [paper]
- π "Rethinking Memory in AI: Taxonomy, Operations, Topics, and Future Directions" [paper]
- "Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs" [paper]
- "Rethinking Agent Design: From Top-Down Workflows to Bottom-Up Skill Evolution" [paper]
- "Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution" [paper]
- "Absolute Zero: Reinforced Self-play Reasoning with Zero Data" [paper]
- "Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks" [paper]
- "DEBATE, TRAIN, EVOLVE: Self-Evolution of Language Model Reasoning" [paper]
- "Self Rewarding Self Improving" [paper]
- "EvolveSearch: An Iterative Self-Evolving Search Agent" [paper]
- "AlphaEvolve: A coding agent for scientific and algorithmic discovery" [paper]
- "Meta-Design Matters:A Self-Design Multi-Agent System" [paper]
- "Darwin GΓΆdel Machine:Open-Ended Evolution of Self-Improving Agents" [paper]
- "SEW: Self-Evolving Agentic Workflows for Automated Code Generation" [paper]
- "Multi-Agent Collaboration via Evolving Orchestration" [paper]
- π "Creativity in LLM-based Multi-Agent Systems: A Survey" [paper]
- βοΈ "Benchmarking LLMsβ Swarm intelligence" [paper]
- "Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems" [paper]
- "Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications" [paper]
- "Towards Multi-Agent Reasoning Systems for Collaborative Expertise Delegation: An Exploratory Design Study" [paper]
- "34 Examples of LLM Applications in Materials Science and Chemistry: Towards Automation, Assistants, Agents, and Accelerated Scientific Discovery" [paper]
- "PiFlow: Principle-aware Scientific Discovery with Multi-Agent Collaboration" [paper]
- "R&D-Agent: Automating Data-Driven AI Solution Building Through LLM-Powered Automated Research, Development, and Evolution" [paper]
- π "From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery" [paper]
- "Towards Artificial Intelligence Research Assistant for Expert-Involved Learning" [paper]
- "MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering" [paper]
- "ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering" [paper]
- "Data-to-Dashboard: Multi-Agent LLM Framework for Insightful Visualization in Enterprise Analytics" [paper]
- "Agentic Feature Augmentation: Unifying Selection and Generation with Teaming, Planning, and Memories" [paper]
- "JARVIS: A Multi-Agent Code Assistant for High-Quality EDA Script Generation" [paper]
- "MLZero: A Multi-Agent System for End-to-end Machine Learning Automation" [paper]
- "Can Agents Fix Agent Issues?" [paper]
- "Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI" [paper]
- "The Real Barrier to LLM Agent Usability is Agentic ROI" [paper]
- π "A Survey on Large Language Model based Human-Agent Systems" [paper]
- π "Vision-Language-Action Models: Concepts, Progress, Applications and Challenges" [paper]
- π "Multi-agent Embodied AI: Advances and Future Directions" [paper]
- "Efficient Agent Training for Computer Use" [paper]
- βοΈ "AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios" [paper]
- "Inference-Time Scaling for Generalist Reward Modeling" [paper]
- "Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead"[paper]
- "Review, Refine, Repeat: Understanding Iterative Decoding of AI Agents with Dynamic Evaluation and Selection"[paper]
- "Dual Engines of Thoughts: A Depth-Breadth Integration Framework for Open-Ended Analysis"[paper]
- π "A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems"[paper]
- "Welcome to the Era of Experience" [paper]
- "SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills"[paper]
- "Exploring Expert Failures Improves LLM Agent Tuning" [paper]
- "Inducing Programmatic Skills for Agentic Tasks" [paper]
- "Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory" [paper]
- "Local Prompt Optimization" [paper]
- "Revisiting Prompt Optimization with Large Reasoning ModelsβA Case Study on Event Extraction" [paper]
- "Iterative Trajectory Exploration for Multimodal Agents" [papaer]
- "FlowReasoner: Reinforcing Query-Level Meta-Agents" [paper]
- "A Self-Improving Coding Agent" [paper]
- "Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models" [paper]
- "ToolRL: Reward is All Tool Learning Needs" [paper]
- "OTC: Optimal Tool Calls via Reinforcement Learning" [paper]
- "LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities" [paper]
- π "Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey" [paper]
- "The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search" [paper]
- "UFO2: The Desktop AgentOS" [paper]
- "AGENTADA: Skill-Adaptive Data Analytics for Tailored Insight Discovery"[paper]
- βοΈ "BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents" [paper]
- "Toward Super Agent System with Hybrid AI Router" [paper] "AgentA/B: Automated and Scalable Web A/B Testing with Interactive LLM Agents" [paper]
- [Apr 2025] "UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents" [paper]
- π "Challenges and Paths Towards AI for Software Engineering"[paper]
- π "Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems"[paper]
- π "Adaptive Human-Agent Teaming: A Review of Empirical Studies from the Process Dynamics Perspective" [paper]
- π "A Survey of AI Agent Protocols" [paper]
π₯: Recommended papers
π: Survey papers
βοΈ: Benchmark papers
- Agent Capabilities
- GenAI Agents Architecture
- GenAI Agents Applications
- GenAI Agents Presentations
