AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,973 results
Agents

Towards Multi-Turn Dialog Systems for Industrial Asset Operations and Maintenance

DGX agent

arXiv:2605.24953v1 Announce Type: new Abstract: Industrial asset operations and maintenance question answering is inherently multi-turn, iterative, and highly dependent on external tool invocation. Ho

agentsarxiv-cs-ai
26 May 2026
Agents
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform

DGX agent

arXiv:2605.23972v1 Announce Type: new Abstract: Large language models achieve strong performance in language generation and knowledge-intensive tasks, yet remain limited in settings requiring causal r

agentsarxiv-cs-ai
26 May 2026
Agents

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

DGX agent

arXiv:2605.23271v1 Announce Type: cross Abstract: The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such deman

agentsarxiv-cs-ai
25 May 2026
Model Releases

Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems

DGX agent

arXiv:2605.23109v1 Announce Type: new Abstract: AI agents increasingly excel at generating, testing, and refining code. However, they fall short on tasks requiring formal guarantees of full coverage t

model-releasesarxiv-cs-ai
25 May 2026
Safety

Understanding Goal Generalisation in Sequential Reinforcement Learning

DGX agent

arXiv:2605.23565v1 Announce Type: cross Abstract: Reinforcement learning agents often exhibit unintended goal-directed behaviour outside their training distribution, but we currently lack a principled

safetyarxiv-cs-ai
25 May 2026
Agents

Can frontier models forecast scientific progress? Mostly no, but here is why. This work looks at 4,760 scientific events across disciplines.…

DGX agent

Can frontier models forecast scientific progress? Mostly no, but here is why. This work looks at 4,760 scientific events across disciplines. Frontier models can identify plausible research directions

agentsdair-ai--x
23 May 2026
Local Ai

Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration

DGX agent

arXiv:2605.22814v1 Announce Type: new Abstract: Exploration is a prerequisite for learning useful behaviors in sparse-reward, long-horizon tasks, particularly within 3D environments. Curiosity-driven

local-aiarxiv-cs-lg
23 May 2026
Model Releases

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

DGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

model-releasesarxiv-cs-cv
22 May 2026
Agents

Kakuna: skills with checklists that only know how to harden your codebase /plan with it then let it /goal for a day, it comes back with same…

DGX agent

Kakuna: skills with checklists that only know how to harden your codebase /plan with it then let it /goal for a day, it comes back with same functionality but all the boring stuff done for you + an au

agentsswyx--x
22 May 2026
Agents

Open weight models running on open source harnesses solve this problem. One of these days, American companies will wake up to the same solut…

DGX agent

Open weight models running on open source harnesses solve this problem. One of these days, American companies will wake up to the same solution we pioneered decades ago that Chinese teams are now runn

agentsyann-lecun--x
22 May 2026
Agents

Psy-Chronicle:A Structured Pipeline for Synthesizing Long-Horizon Campus Psychological Counseling Dialogues

DGX agent

arXiv:2605.22140v1 Announce Type: new Abstract: In recent years, large language models have shown substantial potential in psychological support tasks. However, existing psychological counseling data

agentsarxiv-cs-cl
22 May 2026
Model Releases

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale

DGX agent

arXiv:2605.20744v1 Announce Type: new Abstract: Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby

model-releasesarxiv-cs-lg
21 May 2026
Agents

Towards Resilient and Autonomous Networks: A BlueSky Vision on AI-Native 6G

DGX agent

arXiv:2605.21395v1 Announce Type: cross Abstract: The proliferation of emerging applications, such as autonomous driving and immersive experiences, demands cellular networks that are not only faster,

agentsarxiv-cs-lg
21 May 2026
Model Releases

ZEBRA: Zero-shot Budgeted Resource Allocation for LLM Orchestration

DGX agent

arXiv:2605.20485v1 Announce Type: new Abstract: As autonomous agents increasingly execute end-to-end tasks under fixed monetary budgets, the pressing open question shifts from whether the budget is re

model-releasesarxiv-cs-lg
21 May 2026
Agents

Adaptive Threshold-Driven Continuous Greedy Method for Scalable Submodular Optimization

DGX agent

arXiv:2604.03419v2 Announce Type: replace Abstract: Submodular maximization under matroid constraints is a fundamental problem in combinatorial optimization with applications in sensing, data summariz

agentsarxiv-cs-lg
20 May 2026
Tutorials

Beyond Rational Illusion: Behaviorally Realistic Strategic Classification

DGX agent

arXiv:2605.19674v1 Announce Type: new Abstract: Strategic classification(SC) studies the interaction between decision models and agents who strategically manipulate their features for favorable outcom

tutorialsarxiv-cs-ai
20 May 2026
Safety

Distributional AGI Safety

DGX agent

arXiv:2512.16856v2 Announce Type: replace Abstract: AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an e

safetyarxiv-cs-ai
20 May 2026
Agents

Encryption standards face a reckoning as quantum computing era edges closer

DGX agent

The security landscape is entering uncharted territory as quantum computing moves from theoretical threat to near-term enterprise reality — and the race to post-quantum encryption is one most organiza

agentssiliconangle
20 May 2026
Safety

GAE Falls Short in Imperfect-Information Self-Play Reinforcement Learning

DGX agent

arXiv:2605.19235v1 Announce Type: new Abstract: Competitive multi-agent reinforcement learning in imperfect-information games requires agents to act under partial observability and against adversarial

safetyarxiv-cs-lg
20 May 2026
Model Releases

Google I/O, Gemini Spark, Antigravity

DGX agent

It's hard to find much to write about Google I/O this year because I have a policy of not writing about anything that I can't try out myself, and a lot of the big announcements are 'coming soon'. I ac

model-releasessimon-willison
20 May 2026
Agents

PASC: Pipeline-Aware Conformal Prediction with Joint Coverage Guarantees for Multi-Stage NLP and LLM Pipelines

DGX agent

arXiv:2605.18812v1 Announce Type: cross Abstract: Modern NLP and LLM systems are pipelines: named entity recognition (NER) -> entity disambiguation (NED) -> entity typing, retrieval-augmented generati

agentsarxiv-cs-cl
20 May 2026
Agents

The Wikidata Query Logs Dataset

DGX agent

arXiv:2602.14594v2 Announce Type: replace Abstract: We present the Wikidata Query Logs (WDQL) dataset, a dataset consisting of 335k question-query pairs over the Wikidata knowledge graph. It is over 1

agentsarxiv-cs-cl
20 May 2026
Model Releases

A Machine With Human-Like Memory Systems

DGX agent

arXiv:2204.01611v3 Announce Type: replace Abstract: Inspired by the cognitive science theory, we explicitly model an agent with both semantic and episodic memory systems, and show that it is better th

model-releasesarxiv-cs-ai
19 May 2026
Local Ai

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks

DGX agent

arXiv:2605.18194v1 Announce Type: new Abstract: While Multi-Modal Large Language Models (MLLMs) demonstrate impressive capabilities in general reasoning, their embodied spatial intelligence remains ha

local-aiarxiv-cs-ai
19 May 2026
Research

Counterparty Modeling is Not Strategy: The Limits of LLM Negotiators

DGX agent

arXiv:2605.16575v1 Announce Type: new Abstract: Negotiation requires more than inferring what the other side wants: it requires using that information to make advantageous offers and counteroffers ove

researcharxiv-cs-ai
19 May 2026
Model Releases

From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning

DGX agent

arXiv:2605.17162v1 Announce Type: new Abstract: This paper investigates whether shallow neural network agents can master the card game Schnapsen and challenge a strong search-based baseline, RdeepBot,

model-releasesarxiv-cs-ai
19 May 2026
Agents

Generative AI and Two-Tiered Online Mental Health Communities

DGX agent

arXiv:2605.16279v1 Announce Type: cross Abstract: Online mental health communities (OMHCs) are tiered platforms that connect patients with licensed counselors through public Q&A forums and paid privat

agentsarxiv-cs-ai
19 May 2026
Agents

Genflow Ad Studio: A Compound AI Architecture for Brand-Aligned, Self-Correcting Video Generation

DGX agent

arXiv:2605.16748v1 Announce Type: cross Abstract: Recent advancements in generative video models demonstrate high visual fidelity, yet their integration into enterprise environments is restricted by t

agentsarxiv-cs-ai
19 May 2026
Local Ai

Incentive-Aware Federated Averaging with Performance Guarantees under Strategic Participation

DGX agent

arXiv:2603.20873v2 Announce Type: replace Abstract: Federated learning (FL) is a communication-efficient collaborative learning framework that enables model training across multiple agents with privat

local-aiarxiv-cs-lg
19 May 2026
Agents

Interactive Evaluation Requires a Design Science

DGX agent

arXiv:2605.17829v1 Announce Type: new Abstract: AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, env

agentsarxiv-cs-ai
19 May 2026
Model Releases

Live from Code with Claude London: we're launching self-hosted sandboxes (public beta) and MCP tunnels (research preview) in Claude Managed …

DGX agent

Live from Code with Claude London: we're launching self-hosted sandboxes (public beta) and MCP tunnels (research preview) in Claude Managed Agents. Run agents inside your own perimeter, with your secu

model-releasesboris-cherny--x
19 May 2026
Model Releases

MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning

DGX agent

arXiv:2601.21468v5 Announce Type: replace Abstract: Long-horizon agentic reasoning necessitates effectively compressing growing interaction histories into a limited context window. Most existing memor

model-releasesarxiv-cs-ai
19 May 2026
Agents

Position: Universal Time Series Foundation Models Rest on a Category Error

DGX agent

arXiv:2602.05287v2 Announce Type: replace Abstract: This position paper argues that the pursuit of 'Universal Foundation Models for Time Series' rests on a fundamental category error, mistaking a stru

agentsarxiv-cs-ai
19 May 2026
Model Releases

SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain

DGX agent

arXiv:2605.17946v1 Announce Type: new Abstract: Multimodal large language models are increasingly used as agent backbones that understand multimodal inputs, plan retrieval actions, invoke external too

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

The last six months in LLMs in five minutes

DGX agent

I put together these annotated slides from my five minute lightning talk at PyCon US 2026, using the latest iteration of my annotated presentation tool. # I presented this lightning talk at PyCon US 2

model-releasessimon-willison
19 May 2026
Agents

When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning

DGX agent

arXiv:2605.16312v1 Announce Type: cross Abstract: We study adversarial action masking in self-play reinforcement learning: an attacker selectively removes legal actions from a victim's action set. Unl

agentsarxiv-cs-ai
19 May 2026
Agents

Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding

DGX agent

arXiv:2605.17823v1 Announce Type: cross Abstract: When humans view scenes without a specific task (free-viewing), they initially direct their eye movements toward the scene center and then fixate on p

agentsarxiv-cs-ai
19 May 2026
Agents

Access Timing as Scaffolding: A Reinforcement Learning Approach to GenAI in Education

DGX agent

arXiv:2605.15850v1 Announce Type: cross Abstract: In recent years, generative AI (GenAI) in educational settings has become ubiquitous in students' daily lives, despite its potential to induce over-re

agentsarxiv-cs-ai
18 May 2026
Agents

Lamarckian Inheritance in Dynamic Environments: How Key Variables Affect Evolutionary Dynamics

DGX agent

arXiv:2605.15769v1 Announce Type: cross Abstract: The co-optimization of a robot's body and brain presents a coupled challenge: the morphology constrains which control strategies are effective, while

agentsarxiv-cs-ai
18 May 2026
Hardware

NVIDIA CEO Jensen Huang at Dell Technologies World: “Demand Is Going Parabolic, Utterly Parabolic”

DGX agent

Agentic AI inference at one-tenth the cost per token with NVIDIA Vera Rubin NVL72. Agent sandboxes run 50% faster on NVIDIA Vera than traditional CPUs — while enterprise data queries are up to 3x fast

hardwarenvidia-blog
18 May 2026
Agents

Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data

DGX agent

arXiv:2509.21465v3 Announce Type: replace Abstract: Tabular foundation models are becoming increasingly popular for low-resource tabular problems. These models make up for small training datasets by p

agentsarxiv-cs-lg
18 May 2026
Agents

The best feature of @xai Grok Build right now is how it handles subagents and personas. Most people still treat the model like one very smar…

DGX agent

The best feature of @xai Grok Build right now is how it handles subagents and personas. Most people still treat the model like one very smart intern that has to do everything at once. Grok Build took

agentselon-musk--x
17 May 2026
Agents

Are your benchmarks actually measuring the capability you think they measure? New paper says they probably not. Coined the 'The Evaluation T…

DGX agent

Are your benchmarks actually measuring the capability you think they measure? New paper says they probably not. Coined the 'The Evaluation Trap', it provides a vocabulary for auditing whether your eva

agentsdair-ai--x
16 May 2026
Agents

AI Knows When It's Being Watched: Functional Strategic Action and Contextual Register Modulation in Large Language Models

DGX agent

arXiv:2605.15034v1 Announce Type: cross Abstract: Large language models (LLMs) have been extensively studied from computational and cognitive perspectives, yet their behavior as communicative actors i

agentsarxiv-cs-ai
15 May 2026
Agents

Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning

DGX agent

arXiv:2605.14054v1 Announce Type: new Abstract: Achieving robust perception-reasoning synergy is a central goal for advanced Vision-Language Models (VLMs). Recent advancements have pursued this goal v

agentsarxiv-cs-ai
15 May 2026
Agents

COTCAgent: Preventive Consultation via Probabilistic Chain-of-Thought Completion

DGX agent

arXiv:2605.15016v1 Announce Type: cross Abstract: As large language models empower healthcare, intelligent clinical decision support has developed rapidly. Longitudinal electronic health records (EHR)

agentsarxiv-cs-ai
15 May 2026
Agents

Go in with expectations that Grok Build is still beta, but improving almost every day

DGX agent

Go in with expectations that Grok Build is still beta, but improving almost every day Grok Build is amazing. The early beta just dropped for SuperGrok Heavy users and the first real feedback from deve

agentselon-musk--x
15 May 2026
Agents

Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations

DGX agent

arXiv:2605.14175v1 Announce Type: new Abstract: In long conversations, an LLM can produce a next utterance that sounds plausible but rests on premises the conversation has already abandoned. Context-m

agentsarxiv-cs-ai
15 May 2026
← Previous
1…222223224225226…375
Next →