Skip to content
Β 
Β 

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

54 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

AI Agents Papers

Updated biweekly.

AI Agent

  • AI agents are autonomous software entities that leverage intelligence technologies such as large language models and reinforcement learning to interact with their environment and pursue defined goals.

  • They acquire and generate information through external APIs, sensors, and code execution capabilities, enabling context-aware decision-making and action selection.

  • Through self-learning and feedback loops, they continuously improve their performance, minimizing human intervention while handling complex tasks.

  • By supporting multimodal inputs and coordinating with other agents, they realize diverse cognitive functions such as dialogue, reasoning, and strategic planning.

AI Agent Workflows

June Highlights (Updated 30 June)

Deep Research Agents

  • βš–οΈ "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents" [paper]
  • πŸ“– "From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents" [paper]
  • πŸ“– "Deep Research Agents: A Systematic Examination And Roadmap" [paper]
  • πŸ“– "Towards AI Search Paradigm" [paper]
  • "Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge" [paper]
  • "MMSearch-R1: Incentivizing LMMs to Search" [paper]
  • "Towards Robust Fact-Checking: A Multi-Agent System with Advanced Evidence Retrieval" [paper]

Data Science Agents

  • [Jun 2025] "AUTOMIND: Adaptive Knowledgeable Agent for Automated Data Science" [paper]
  • πŸ“– [Jun 2025] "Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents" [paper]
  • [Jun 2025] "SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation" [paper]
  • [Jun 2025] "SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications" [paper]
  • [Jun 2025] "Towards Community-Driven Agents for Machine Learning Engineering" [paper]
  • [Jun 2025] "MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement" [paper]

Business Operation Agents

  • "Oversight Structures for Agentic AI in Public-Sector Organizations" [paper]
  • βš–οΈ "AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance" [paper]
  • πŸ“– "Application-Driven Value Alignment in Agentic AI Systems: Survey and Perspectives" [paper]
  • "Intelligent Design 4.0: Paradigm Evolution Toward the Agentic AI Era" [paper]
  • "Improved LLM Agents for Financial Document Question Answering" [paper]
  • βš–οΈ "ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering" [paper]
  • "Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine" [paper]
  • "SV-LLM: An Agentic Approach for SoC Security Verification using Large Language Models" [paper]
  • βš–οΈ "SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents" [paper]
  • "Intelligent Design 4.0: Paradigm Evolution Toward the Agentic AI Era" [paper]
  • "Managing Complex Failure Analysis Workflows with LLM-based Reasoning and Acting Agents" [paper]
  • "AgenticControl: An Automated Control Design Framework Using Large Language Models" [paper]
  • πŸ“– "A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools" [paper]

May Highlights (Updated 31 May)

Inference Time Computing

  • πŸ“– "A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law" [paper]
  • πŸ“– "Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models" [paper]

Tool Integrated Reasoning

  • "Table-R1: Inference-Time Scaling for Table Reasoning" [paper]
  • "Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning" [paper]
  • "Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning" [paper]
  • "Agent RL Scaling Law: Spontaneous Code Execution for Mathematical Problem Solving" [paper]
  • "Reinforced Internal-External Knowledge Synergistic Reasoning for Efficient Adaptive Search Agent" [paper]
  • "An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents" [paper]
  • "Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning" [paper]
  • "MIRROR: Multi-agent Intra- and Inter-Reflection for Optimized Reasoning in Tool Learning" [paper]
  • "EvolveSearch: An Iterative Self-Evolving Search Agent" [paper]
  • "VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection" [paper]
  • "Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning" [paper]

Self-Improvement & Self-Evolution

Metric & Reward

  • "RM-R1: Reward Modeling as Reasoning" [paper]
  • "Reward Reasoning Model" [paper]
  • "R3: Robust Rubric-Agnostic Reward Models" [paper]
  • "AutoLibra: Agent Metric Induction from Open-Ended Feedback" [paper]

Memory

  • "MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models (Short Version)" [paper]
  • "MemEngine: A Unified and Modular Library for Developing Advanced Memory of LLM-based Agents" [paper]
  • "MARK: Memory Augmented Refinement of Knowledge" [paper]
  • πŸ“– "Rethinking Memory in AI: Taxonomy, Operations, Topics, and Future Directions" [paper]

Skills

  • "Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs" [paper]
  • "Rethinking Agent Design: From Top-Down Workflows to Bottom-Up Skill Evolution" [paper]
  • "Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution" [paper]

Reasoning Model

  • "Absolute Zero: Reinforced Self-play Reasoning with Zero Data" [paper]
  • "Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks" [paper]
  • "DEBATE, TRAIN, EVOLVE: Self-Evolution of Language Model Reasoning" [paper]
  • "Self Rewarding Self Improving" [paper]
  • "EvolveSearch: An Iterative Self-Evolving Search Agent" [paper]

(Multi) Agent Architecture

  • "AlphaEvolve: A coding agent for scientific and algorithmic discovery" [paper]
  • "Meta-Design Matters:A Self-Design Multi-Agent System" [paper]
  • "Darwin GΓΆdel Machine:Open-Ended Evolution of Self-Improving Agents" [paper]
  • "SEW: Self-Evolving Agentic Workflows for Automated Code Generation" [paper]
  • "Multi-Agent Collaboration via Evolving Orchestration" [paper]

Multi-Agent

  • πŸ“– "Creativity in LLM-based Multi-Agent Systems: A Survey" [paper]
  • βš–οΈ "Benchmarking LLMs’ Swarm intelligence" [paper]
  • "Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems" [paper]
  • "Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications" [paper]
  • "Towards Multi-Agent Reasoning Systems for Collaborative Expertise Delegation: An Exploratory Design Study" [paper]

Real-World Application of AI Agents

Researcher

  • "34 Examples of LLM Applications in Materials Science and Chemistry: Towards Automation, Assistants, Agents, and Accelerated Scientific Discovery" [paper]
  • "PiFlow: Principle-aware Scientific Discovery with Multi-Agent Collaboration" [paper]
  • "R&D-Agent: Automating Data-Driven AI Solution Building Through LLM-Powered Automated Research, Development, and Evolution" [paper]
  • πŸ“– "From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery" [paper]
  • "Towards Artificial Intelligence Research Assistant for Expert-Involved Learning" [paper]

Data Scientist

  • "MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering" [paper]
  • "ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering" [paper]
  • "Data-to-Dashboard: Multi-Agent LLM Framework for Insightful Visualization in Enterprise Analytics" [paper]
  • "Agentic Feature Augmentation: Unifying Selection and Generation with Teaming, Planning, and Memories" [paper]
  • "JARVIS: A Multi-Agent Code Assistant for High-Quality EDA Script Generation" [paper]
  • "MLZero: A Multi-Agent System for End-to-end Machine Learning Automation" [paper]

Software Engineer

  • "Can Agents Fix Agent Issues?" [paper]
  • "Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI" [paper]

Others

  • "The Real Barrier to LLM Agent Usability is Agentic ROI" [paper]
  • πŸ“– "A Survey on Large Language Model based Human-Agent Systems" [paper]
  • πŸ“– "Vision-Language-Action Models: Concepts, Progress, Applications and Challenges" [paper]
  • πŸ“– "Multi-agent Embodied AI: Advances and Future Directions" [paper]
  • "Efficient Agent Training for Computer Use" [paper]
  • βš–οΈ "AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios" [paper]

April Highlights

Inference Time Computing

  • "Inference-Time Scaling for Generalist Reward Modeling" [paper]
  • "Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead"[paper]
  • "Review, Refine, Repeat: Understanding Iterative Decoding of AI Agents with Dynamic Evaluation and Selection"[paper]
  • "Dual Engines of Thoughts: A Depth-Breadth Integration Framework for Open-Ended Analysis"[paper]
  • πŸ“– "A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems"[paper]

Self-Experience-Driven Agents

  • "Welcome to the Era of Experience" [paper]
  • "SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills"[paper]
  • "Exploring Expert Failures Improves LLM Agent Tuning" [paper]
  • "Inducing Programmatic Skills for Agentic Tasks" [paper]
  • "Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory" [paper]
  • "Local Prompt Optimization" [paper]
  • "Revisiting Prompt Optimization with Large Reasoning Modelsβ€”A Case Study on Event Extraction" [paper]
  • "Iterative Trajectory Exploration for Multimodal Agents" [papaer]

Meta Agents

  • "FlowReasoner: Reinforcing Query-Level Meta-Agents" [paper]
  • "A Self-Improving Coding Agent" [paper]
  • "Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models" [paper]

Reinforcement Learning Applications for AI Agents

  • "ToolRL: Reward is All Tool Learning Needs" [paper]
  • "OTC: Optimal Tool Calls via Reinforcement Learning" [paper]
  • "LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities" [paper]
  • πŸ“– "Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey" [paper]

Real-World Application of AI Agents

  • "The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search" [paper]
  • "UFO2: The Desktop AgentOS" [paper]
  • "AGENTADA: Skill-Adaptive Data Analytics for Tailored Insight Discovery"[paper]
  • βš–οΈ "BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents" [paper]
  • "Toward Super Agent System with Hybrid AI Router" [paper] "AgentA/B: Automated and Scalable Web A/B Testing with Interactive LLM Agents" [paper]
  • [Apr 2025] "UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents" [paper]
  • πŸ“– "Challenges and Paths Towards AI for Software Engineering"[paper]

Survey

  • πŸ“– "Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems"[paper]
  • πŸ“– "Adaptive Human-Agent Teaming: A Review of Empirical Studies from the Process Dynamics Perspective" [paper]
  • πŸ“– "A Survey of AI Agent Protocols" [paper]

Paper Categories

πŸ”₯: Recommended papers
πŸ“–: Survey papers
βš–οΈ: Benchmark papers

References

About

A collection of AI Agents papers (Updated biweekly)

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors