AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,979 results
Model Releases

ProSPy: A Profiling-Driven SQL-Python Agentic Framework for Enterprise Text-to-SQL

DGX agent

arXiv:2606.05836v1 Announce Type: new Abstract: Large language models have substantially advanced Text-to-SQL systems, yet applying them to enterprise-scale databases remains challenging. Real-world d

model-releasesarxiv-cs-cl
5 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

The Meta hack shows there’s more to AI security than Mythos

DGX agent

On June 5, 404 Media reported that attackers had been using Meta’s AI customer support agent to steal Instagram accounts. Their approach was simple: They asked the agent to link the accounts to email

agentsmit-tech-review
5 Jun 2026
Agents

Consensus is Strategically Insufficient: Reasoning-Trace Disagreement as a Knowledge-Representation Signal

DGX agent

arXiv:2606.04223v1 Announce Type: new Abstract: Multi-agent systems are commonly designed to reduce disagreement through voting, consensus protocols, debate, or fault-tolerant aggregation. We argue th

agentsarxiv-cs-ai
4 Jun 2026
Model Releases

HighTide: An Agent-Curated Open-Source VLSI Benchmark Suite

DGX agent

arXiv:2606.04126v1 Announce Type: cross Abstract: We introduce HighTide, an evolving AI-assisted benchmark suite. Specifically, the contributions are: (i) a diverse open-source suite spanning multiple

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Impostor: An Agent-Curated Benchmark for Realistic AIGC Manipulation Localization

DGX agent

arXiv:2606.04545v1 Announce Type: new Abstract: Recent advances in generative image editing have improved the realism and controllability of localized image manipulation, raising new challenges for im

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

LifeSide: Benchmarking Agents as Lifelong Digital Companions

DGX agent

arXiv:2606.04660v1 Announce Type: new Abstract: Lifelong digital companions must integrate cross-session cues, continually update their understanding of users, and adapt to shifting privacy boundaries

model-releasesarxiv-cs-cl
4 Jun 2026
Safety

Neetyabhas: A Framework for Uncertainty-Aware Public Policy Optimization in Rational Agent-Based Models

DGX agent

arXiv:2606.04562v1 Announce Type: new Abstract: Purpose The WHO's COVID-19 non-pharmaceutical interventions (e.g., lockdowns, vaccinations) effectively curb transmission but impose heavy economic stra

safetyarxiv-cs-ai
4 Jun 2026
Safety

Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents

DGX agent

arXiv:2510.13704v2 Announce Type: replace-cross Abstract: Recent works have proposed accelerating the wall-clock training time of actor-critic methods via the use of large-scale environment paralleliz

safetyarxiv-cs-ai
4 Jun 2026
Agents

Sometimes you need to start over. But that decision is hard. @llama_index had to make that call: they built one of the most popular AI frame…

DGX agent

Sometimes you need to start over. But that decision is hard. @llama_index had to make that call: they built one of the most popular AI frameworks in the world, but saw the agent harness and frontier l

agentsjerry-liu--x
4 Jun 2026
Model Releases

Towards Efficient and Evidence-grounded Mobility Prediction with LLM-Driven Agent

DGX agent

arXiv:2606.05130v1 Announce Type: cross Abstract: Individual-level mobility prediction is central to urban simulation, transportation planning, and policy analysis. Supervised sequence models achieve

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

AI Agents Enable Adaptive Computer Worms

DGX agent

arXiv:2606.03811v1 Announce Type: cross Abstract: A computer worm is malware that spreads on a network by replicating itself from one machine to another. Traditional worms, like WannaCry, exploited pr

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents

DGX agent

arXiv:2606.03829v1 Announce Type: new Abstract: Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and a

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach

DGX agent

arXiv:2510.23216v4 Announce Type: replace Abstract: While several high profile video games have served as testbeds for Deep Reinforcement Learning (DRL), this technique has rarely been employed by the

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

DGX agent

arXiv:2606.03168v1 Announce Type: new Abstract: While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

DGX agent

arXiv:2606.03303v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong informal mathematical reasoning but struggle to generate mechanically verifiable proofs in formal languages

model-releasesarxiv-cs-ai
3 Jun 2026
Hardware

Scaling Multi Agent Reinforcement Learning for Underwater Acoustic Tracking via Autonomous Vehicles

DGX agent

arXiv:2505.08222v3 Announce Type: replace-cross Abstract: Autonomous vehicles (AVs) offer a cost-effective solution for scientific missions such as underwater tracking. Reinforcement learning (RL) has

hardwarearxiv-cs-ai
3 Jun 2026
Model Releases

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill

DGX agent

arXiv:2606.03980v1 Announce Type: cross Abstract: Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) p

model-releasesarxiv-cs-cl
3 Jun 2026
Agents

This SkillOpt paper from Microsoft is a must-read! (bookmark it) I was a bit skeptical of the results reported in the paper when I shared it…

DGX agent

This SkillOpt paper from Microsoft is a must-read! (bookmark it) I was a bit skeptical of the results reported in the paper when I shared it a few days ago. However, I managed to integrate it into my

agentsdair-ai--x
3 Jun 2026
Model Releases

3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

DGX agent

arXiv:2606.01057v1 Announce Type: cross Abstract: Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neur

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models

DGX agent

arXiv:2606.02459v1 Announce Type: new Abstract: Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is d

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design

DGX agent

arXiv:2606.02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical f

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework

DGX agent

arXiv:2602.18008v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise in constructing mechanistic models from data. However, existing evaluations largely focus on s

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

DGX agent

arXiv:2606.01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed samp

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes

DGX agent

arXiv:2606.02215v1 Announce Type: new Abstract: Large Language Model (LLM)-augmented Community Notes offer a scalable path for timely, evidence-grounded correction of health misinformation on social p

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents

DGX agent

arXiv:2606.02568v1 Announce Type: new Abstract: Clinical practice is not the selection of an answer from enumerated options: a physician gathers heterogeneous information incrementally and commits to

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

Dive into Ambiguity: A*-Inspired Multi-Agents Commonsense Obfuscation Attack on LLM Prompts

DGX agent

arXiv:2606.01441v1 Announce Type: new Abstract: Large language models (LLMs) excel in reasoning and knowledge-intensive tasks but remain vulnerable to prompt-level adversarial attacks that preserve in

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering

DGX agent

arXiv:2606.01434v1 Announce Type: new Abstract: Drug-information question answering is a high-stakes setting where hallucinated facts can mislead clinical decision-making and the provenance of each ci

model-releasesarxiv-cs-cl
2 Jun 2026
Agents

Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity

DGX agent

arXiv:2606.01637v1 Announce Type: cross Abstract: Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a m

agentsarxiv-cs-ai
2 Jun 2026
Agents

Four insights you might have missed from theCUBE’s coverage of KB4-CON

DGX agent

Artificial intelligence agents are turning agent risk management into a frontline security priority. As digital workers begin touching email, financial systems, collaboration tools and business workfl

agentssiliconangle
2 Jun 2026
Local Ai

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs

DGX agent

arXiv:2606.00104v1 Announce Type: cross Abstract: Foundation models are increasingly used to drive autonomous systems, yet existing approaches either keep the model in a tight control loop, raising la

local-aiarxiv-cs-ai
2 Jun 2026
Model Releases

SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes

DGX agent

arXiv:2606.01912v1 Announce Type: new Abstract: Smart homes are evolving toward complex state-dependent living environments, requiring Large Language Models (LLMs) to reason over user intent, preferen

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

DGX agent

arXiv:2606.01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existi

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination

DGX agent

arXiv:2606.01276v1 Announce Type: new Abstract: Large language model (LLM)-based machine translation has advanced cross-cultural communication, yet it still struggles with culture-loaded words (CLWs)

model-releasesarxiv-cs-cl
2 Jun 2026
Agents

CoMem: Context Management with A Decoupled Long-Context Model

DGX agent

arXiv:2605.30842v1 Announce Type: new Abstract: Context management enables agentic models to solve long-horizon tasks through iterative summarization of previous interaction histories. However, this p

agentsarxiv-cs-lg
1 Jun 2026
Safety

Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents

DGX agent

arXiv:2605.31354v1 Announce Type: new Abstract: Modular visual reasoning systems increasingly rely on shared working memory for multi-step collaboration, yet the failure dynamics of intermediate state

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

DGX agent

arXiv:2605.31148v1 Announce Type: cross Abstract: Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into ac

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation

DGX agent

arXiv:2605.30916v1 Announce Type: new Abstract: AI benchmarks have well-documented limitations, with prior work examining contamination, saturation, and construct underspecification. Aggregation has r

model-releasesarxiv-cs-lg
1 Jun 2026
Research

Analyzing Persona Effects in Generated Explanations from Multimodal LLM Agents in Urban Perception

DGX agent

arXiv:2605.29064v1 Announce Type: new Abstract: We study how persona prompting shapes language generated by multimodal large language models in an urban perception setting. Using 59,808 annotations fr

researcharxiv-cs-cl
29 May 2026
Model Releases

AutoSizer: Automatic Sizing of Analog and Mixed-Signal Circuits via Large Language Model (LLM) Agents

DGX agent

arXiv:2602.02849v2 Announce Type: replace Abstract: The design of Analog and Mixed-Signal (AMS) integrated circuits remains heavily reliant on expert knowledge, with transistor sizing a major bottlene

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation

DGX agent

arXiv:2605.30090v1 Announce Type: new Abstract: Long-form video generation is rapidly moving from short, single-scene synthesis toward minute-long, multi-shot creation with narrative structure, cinema

model-releasesarxiv-cs-cl
29 May 2026
Safety

Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling

DGX agent

arXiv:2605.29262v1 Announce Type: new Abstract: The Dynamic Flexible Job Shop Scheduling Problem (DFJSP) necessitates a trade-off between instant reaction to stochastic disturbances and global optimiz

safetyarxiv-cs-ai
29 May 2026
Model Releases

Learning Design Skills as Memory Policies for Agentic Photonic Inverse Design

DGX agent

arXiv:2605.29421v1 Announce Type: new Abstract: Photonic crystal fiber (PCF) inverse design remains challenging because candidate geometries must satisfy coupled optical targets under expensive electr

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection

DGX agent

arXiv:2605.30274v1 Announce Type: cross Abstract: Document-level translation remains one of the most challenging tasks for large language models, which are constrained by limited context windows that

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software

DGX agent

arXiv:2605.30353v1 Announce Type: new Abstract: Are AI agents tools, co-authors, or researchers? We present a quantified case study (N=1): a physicist supervising an AI coding agent (Claude Code, Sonn

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Small Agent Group is the Future of Digital Health

DGX agent

arXiv:2602.08013v2 Announce Type: replace Abstract: The rapid adoption of large language models (LLMs) in digital health has been driven by a 'scaling-first' philosophy, i.e., the assumption that clin

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data

DGX agent

arXiv:2605.29483v1 Announce Type: new Abstract: Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-spe

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

DGX agent

arXiv:2605.28056v1 Announce Type: new Abstract: Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict

DGX agent

arXiv:2605.28369v1 Announce Type: new Abstract: E-commerce platforms have begun recruiting crowdsourced jurors to adjudicate massive volumes of transaction disputes. Unlike formal legal judgment, E-co

model-releasesarxiv-cs-ai
28 May 2026
← Previous
1…182183184185186…375
Next →