AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
28 May 2026

Personalized Observation Normalization for Federated Reinforcement Learning in Simulation Environments with Heterogeneity

Model ReleasesDGX agent

arXiv:2605.27385v1 Announce Type: cross Abstract: Federated reinforcement learning (FedRL) enables multiple agents to collaboratively train a global policy without sharing raw data, making it ideal fo

You Live More Than Once: Towards Hierarchical Skill Meta-Evolving

Model ReleasesDGX agent

arXiv:2605.28390v1 Announce Type: new Abstract: Test-time skill evolving is regarded as a new paradigm for enhancing deployed agentic systems. Existing works mainly focus on hard-coded skill evolving

27 May 2026

Constrained Meta Reinforcement Learning with Provable Test-Time Safety

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

arXiv:2601.21845v2 Announce Type: replace Abstract: Meta reinforcement learning (RL) allows agents to leverage experience across a distribution of tasks on which the agent can train at will, enabling

Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control

AgentsDGX agent

arXiv:2605.26754v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) increasingly underpins high-stakes applications, yet remains vulnerable to Confundo-style poisoning where adversa

FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentation

AgentsDGX agent

arXiv:2605.27178v1 Announce Type: cross Abstract: We address the challenging task of 3D object segmentation in complex scene point clouds without relying on any scene-level human annotations during tr

Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations

Model ReleasesDGX agent

arXiv:2605.26874v1 Announce Type: cross Abstract: LLM-based agents for industrial asset operations show limited accuracy when reasoning over flat document stores. AssetOpsBench (KDD 2026) establishes

Microsoft and Dell believe the answer to rising cloud token costs sits on every employee’s desk

Local AiDGX agent

Enterprises are adopting an agentic AI PC strategy as Copilot+ PCs shift AI workloads from expensive cloud inference to secure, high-performance on-device processing. As agentic AI moves from experime

The MiniMax M2 series was one of the most widely used open-weight LLM series earlier this year. Now, we got a technical report with some int…

Model ReleasesDGX agent

The MiniMax M2 series was one of the most widely used open-weight LLM series earlier this year. Now, we got a technical report with some interesting tidbits. I summarized some of them below: 1. Full a

26 May 2026

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery

AgentsDGX agent

arXiv:2604.05550v2 Announce Type: replace Abstract: Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-T

EfficientGraph-RAG: Structured Retrieval-State Management for Cross-Task Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2605.25379v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has become the standard way to ground large language models in external knowledge, but many systems still organize

Hypothesis Generation and Inductive Inference in Children and Language Models

ApplicationsDGX agent

arXiv:2605.24528v1 Announce Type: new Abstract: Real world decision-making requires constructing mental models under uncertainty over evidence, over the underlying causal rules, and over the state of

Meta-Engineering Harnesses for AI-Native Software Production: A Contract-Driven Adversarial Verification Architecture with Early Deployment Report

AgentsDGX agent

arXiv:2605.25665v1 Announce Type: cross Abstract: AI-native software development is often evaluated at the level of individual models, prompts, or generated artifacts. This framing is insufficient for

Multi-Persona Debate System for Automated Scientific Hypothesis Generation

AgentsDGX agent

arXiv:2605.23917v1 Announce Type: new Abstract: Modern scientific discovery is bottlenecked not by data scarcity, but by the inability to synthesize fragmented knowledge into actionable hypotheses. Th

PCGRLLM: Large Language Model-Driven Reward Design for Procedural Content Generation Reinforcement Learning

AgentsDGX agent

arXiv:2502.10906v2 Announce Type: replace Abstract: Reward design plays a pivotal role in the training of game AIs, requiring substantial domain-specific knowledge and human effort. In recent years, s

Towards Multi-Turn Dialog Systems for Industrial Asset Operations and Maintenance

AgentsDGX agent

arXiv:2605.24953v1 Announce Type: new Abstract: Industrial asset operations and maintenance question answering is inherently multi-turn, iterative, and highly dependent on external tool invocation. Ho

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform

AgentsDGX agent

arXiv:2605.23972v1 Announce Type: new Abstract: Large language models achieve strong performance in language generation and knowledge-intensive tasks, yet remain limited in settings requiring causal r

25 May 2026

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

AgentsDGX agent

arXiv:2605.23271v1 Announce Type: cross Abstract: The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such deman

Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems

Model ReleasesDGX agent

arXiv:2605.23109v1 Announce Type: new Abstract: AI agents increasingly excel at generating, testing, and refining code. However, they fall short on tasks requiring formal guarantees of full coverage t

Understanding Goal Generalisation in Sequential Reinforcement Learning

SafetyDGX agent

arXiv:2605.23565v1 Announce Type: cross Abstract: Reinforcement learning agents often exhibit unintended goal-directed behaviour outside their training distribution, but we currently lack a principled

23 May 2026

Can frontier models forecast scientific progress? Mostly no, but here is why. This work looks at 4,760 scientific events across disciplines.…

AgentsDGX agent

Can frontier models forecast scientific progress? Mostly no, but here is why. This work looks at 4,760 scientific events across disciplines. Frontier models can identify plausible research directions

Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration

Local AiDGX agent

arXiv:2605.22814v1 Announce Type: new Abstract: Exploration is a prerequisite for learning useful behaviors in sparse-reward, long-horizon tasks, particularly within 3D environments. Curiosity-driven

22 May 2026

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

Model ReleasesDGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

Kakuna: skills with checklists that only know how to harden your codebase /plan with it then let it /goal for a day, it comes back with same…

AgentsDGX agent

Kakuna: skills with checklists that only know how to harden your codebase /plan with it then let it /goal for a day, it comes back with same functionality but all the boring stuff done for you + an au

Open weight models running on open source harnesses solve this problem. One of these days, American companies will wake up to the same solut…

AgentsDGX agent

Open weight models running on open source harnesses solve this problem. One of these days, American companies will wake up to the same solution we pioneered decades ago that Chinese teams are now runn

Psy-Chronicle:A Structured Pipeline for Synthesizing Long-Horizon Campus Psychological Counseling Dialogues

AgentsDGX agent

arXiv:2605.22140v1 Announce Type: new Abstract: In recent years, large language models have shown substantial potential in psychological support tasks. However, existing psychological counseling data

21 May 2026

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale

Model ReleasesDGX agent

arXiv:2605.20744v1 Announce Type: new Abstract: Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby

Towards Resilient and Autonomous Networks: A BlueSky Vision on AI-Native 6G

AgentsDGX agent

arXiv:2605.21395v1 Announce Type: cross Abstract: The proliferation of emerging applications, such as autonomous driving and immersive experiences, demands cellular networks that are not only faster,

ZEBRA: Zero-shot Budgeted Resource Allocation for LLM Orchestration

Model ReleasesDGX agent

arXiv:2605.20485v1 Announce Type: new Abstract: As autonomous agents increasingly execute end-to-end tasks under fixed monetary budgets, the pressing open question shifts from whether the budget is re

20 May 2026

Adaptive Threshold-Driven Continuous Greedy Method for Scalable Submodular Optimization

AgentsDGX agent

arXiv:2604.03419v2 Announce Type: replace Abstract: Submodular maximization under matroid constraints is a fundamental problem in combinatorial optimization with applications in sensing, data summariz

Beyond Rational Illusion: Behaviorally Realistic Strategic Classification

TutorialsDGX agent

arXiv:2605.19674v1 Announce Type: new Abstract: Strategic classification(SC) studies the interaction between decision models and agents who strategically manipulate their features for favorable outcom

Distributional AGI Safety

SafetyDGX agent

arXiv:2512.16856v2 Announce Type: replace Abstract: AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an e

Encryption standards face a reckoning as quantum computing era edges closer

AgentsDGX agent

The security landscape is entering uncharted territory as quantum computing moves from theoretical threat to near-term enterprise reality — and the race to post-quantum encryption is one most organiza

GAE Falls Short in Imperfect-Information Self-Play Reinforcement Learning

SafetyDGX agent

arXiv:2605.19235v1 Announce Type: new Abstract: Competitive multi-agent reinforcement learning in imperfect-information games requires agents to act under partial observability and against adversarial

Google I/O, Gemini Spark, Antigravity

Model ReleasesDGX agent

It's hard to find much to write about Google I/O this year because I have a policy of not writing about anything that I can't try out myself, and a lot of the big announcements are 'coming soon'. I ac

PASC: Pipeline-Aware Conformal Prediction with Joint Coverage Guarantees for Multi-Stage NLP and LLM Pipelines

AgentsDGX agent

arXiv:2605.18812v1 Announce Type: cross Abstract: Modern NLP and LLM systems are pipelines: named entity recognition (NER) -> entity disambiguation (NED) -> entity typing, retrieval-augmented generati

The Wikidata Query Logs Dataset

AgentsDGX agent

arXiv:2602.14594v2 Announce Type: replace Abstract: We present the Wikidata Query Logs (WDQL) dataset, a dataset consisting of 335k question-query pairs over the Wikidata knowledge graph. It is over 1

19 May 2026

A Machine With Human-Like Memory Systems

Model ReleasesDGX agent

arXiv:2204.01611v3 Announce Type: replace Abstract: Inspired by the cognitive science theory, we explicitly model an agent with both semantic and episodic memory systems, and show that it is better th

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks

Local AiDGX agent

arXiv:2605.18194v1 Announce Type: new Abstract: While Multi-Modal Large Language Models (MLLMs) demonstrate impressive capabilities in general reasoning, their embodied spatial intelligence remains ha

Counterparty Modeling is Not Strategy: The Limits of LLM Negotiators

ResearchDGX agent

arXiv:2605.16575v1 Announce Type: new Abstract: Negotiation requires more than inferring what the other side wants: it requires using that information to make advantageous offers and counteroffers ove

From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.17162v1 Announce Type: new Abstract: This paper investigates whether shallow neural network agents can master the card game Schnapsen and challenge a strong search-based baseline, RdeepBot,

Generative AI and Two-Tiered Online Mental Health Communities

AgentsDGX agent

arXiv:2605.16279v1 Announce Type: cross Abstract: Online mental health communities (OMHCs) are tiered platforms that connect patients with licensed counselors through public Q&A forums and paid privat

Genflow Ad Studio: A Compound AI Architecture for Brand-Aligned, Self-Correcting Video Generation

AgentsDGX agent

arXiv:2605.16748v1 Announce Type: cross Abstract: Recent advancements in generative video models demonstrate high visual fidelity, yet their integration into enterprise environments is restricted by t

Incentive-Aware Federated Averaging with Performance Guarantees under Strategic Participation

Local AiDGX agent

arXiv:2603.20873v2 Announce Type: replace Abstract: Federated learning (FL) is a communication-efficient collaborative learning framework that enables model training across multiple agents with privat

Interactive Evaluation Requires a Design Science

AgentsDGX agent

arXiv:2605.17829v1 Announce Type: new Abstract: AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, env

Live from Code with Claude London: we're launching self-hosted sandboxes (public beta) and MCP tunnels (research preview) in Claude Managed …

Model ReleasesDGX agent

Live from Code with Claude London: we're launching self-hosted sandboxes (public beta) and MCP tunnels (research preview) in Claude Managed Agents. Run agents inside your own perimeter, with your secu

MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning

Model ReleasesDGX agent

arXiv:2601.21468v5 Announce Type: replace Abstract: Long-horizon agentic reasoning necessitates effectively compressing growing interaction histories into a limited context window. Most existing memor

Position: Universal Time Series Foundation Models Rest on a Category Error

AgentsDGX agent

arXiv:2602.05287v2 Announce Type: replace Abstract: This position paper argues that the pursuit of 'Universal Foundation Models for Time Series' rests on a fundamental category error, mistaking a stru

SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain

Model ReleasesDGX agent

arXiv:2605.17946v1 Announce Type: new Abstract: Multimodal large language models are increasingly used as agent backbones that understand multimodal inputs, plan retrieval actions, invoke external too

The last six months in LLMs in five minutes

Model ReleasesDGX agent

I put together these annotated slides from my five minute lightning talk at PyCon US 2026, using the latest iteration of my annotated presentation tool. # I presented this lightning talk at PyCon US 2

When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning

AgentsDGX agent

arXiv:2605.16312v1 Announce Type: cross Abstract: We study adversarial action masking in self-play reinforcement learning: an attacker selectively removes legal actions from a victim's action set. Unl

Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding

AgentsDGX agent

arXiv:2605.17823v1 Announce Type: cross Abstract: When humans view scenes without a specific task (free-viewing), they initially direct their eye movements toward the scene center and then fixate on p

18 May 2026

Access Timing as Scaffolding: A Reinforcement Learning Approach to GenAI in Education

AgentsDGX agent

arXiv:2605.15850v1 Announce Type: cross Abstract: In recent years, generative AI (GenAI) in educational settings has become ubiquitous in students' daily lives, despite its potential to induce over-re

Lamarckian Inheritance in Dynamic Environments: How Key Variables Affect Evolutionary Dynamics

AgentsDGX agent

arXiv:2605.15769v1 Announce Type: cross Abstract: The co-optimization of a robot's body and brain presents a coupled challenge: the morphology constrains which control strategies are effective, while

NVIDIA CEO Jensen Huang at Dell Technologies World: “Demand Is Going Parabolic, Utterly Parabolic”

HardwareDGX agent

Agentic AI inference at one-tenth the cost per token with NVIDIA Vera Rubin NVL72. Agent sandboxes run 50% faster on NVIDIA Vera than traditional CPUs — while enterprise data queries are up to 3x fast

Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data

AgentsDGX agent

arXiv:2509.21465v3 Announce Type: replace Abstract: Tabular foundation models are becoming increasingly popular for low-resource tabular problems. These models make up for small training datasets by p

17 May 2026

The best feature of @xai Grok Build right now is how it handles subagents and personas. Most people still treat the model like one very smar…

AgentsDGX agent

The best feature of @xai Grok Build right now is how it handles subagents and personas. Most people still treat the model like one very smart intern that has to do everything at once. Grok Build took

16 May 2026

Are your benchmarks actually measuring the capability you think they measure? New paper says they probably not. Coined the 'The Evaluation T…

AgentsDGX agent

Are your benchmarks actually measuring the capability you think they measure? New paper says they probably not. Coined the 'The Evaluation Trap', it provides a vocabulary for auditing whether your eva

15 May 2026

AI Knows When It's Being Watched: Functional Strategic Action and Contextual Register Modulation in Large Language Models

AgentsDGX agent

arXiv:2605.15034v1 Announce Type: cross Abstract: Large language models (LLMs) have been extensively studied from computational and cognitive perspectives, yet their behavior as communicative actors i

Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning

AgentsDGX agent

arXiv:2605.14054v1 Announce Type: new Abstract: Achieving robust perception-reasoning synergy is a central goal for advanced Vision-Language Models (VLMs). Recent advancements have pursued this goal v

COTCAgent: Preventive Consultation via Probabilistic Chain-of-Thought Completion

AgentsDGX agent

arXiv:2605.15016v1 Announce Type: cross Abstract: As large language models empower healthcare, intelligent clinical decision support has developed rapidly. Longitudinal electronic health records (EHR)

← Previous
1…177178179180181…300
Next →