AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,608 results
6 May 2026

we're continuing to see clear examples where a model's harness is a major determinant of overall performance. with the same model, running o…

Model ReleasesDGX agent

we're continuing to see clear examples where a model's harness is a major determinant of overall performance. with the same model, running on same task, it's easy to observe very different scores depe

5 May 2026

Accelerating battery research with an AI interface between FINALES and Kadi4Mat

Model ReleasesDGX agent

arXiv:2605.00909v1 Announce Type: cross Abstract: The time-consuming formation process critically impacts the longevity of sodium-ion coin cells and End Of Life (EOL) performance. This study aims to o

Accurate Legal Reasoning at Scale: Neuro-Symbolic Offloading and Structural Auditability for Robust Legal Adjudication

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2605.02472v1 Announce Type: new Abstract: Legal texts often contain computational legal clauses--provisions whose understanding requires complex logic. While frontier Large Reasoning Models (LRM

Active Reasoning Vision-Language Models via Sequential Experimental Design

ResearchDGX agent

arXiv:2605.01345v1 Announce Type: new Abstract: Visual perception in modern Vision-Language Models (VLMs) is constrained by a fundamental perceptual bandwidth bottleneck: a broad field of view inevita

ARGUS: Policy-Adaptive Ad Governance via Evolving Reinforcement with Adversarial Umpiring

SafetyDGX agent

arXiv:2605.02200v1 Announce Type: new Abstract: Online advertising governance faces significant challenges due to the non-stationary nature of regulatory policies, where emerging mandates (e.g., restr

Artificial intelligence language technologies in multilingual healthcare: Grand challenges ahead

SafetyDGX agent

arXiv:2605.01441v1 Announce Type: new Abstract: AI language technologies (AILTs), increasingly enabled by large language models (LLMs), are becoming embedded in multilingual healthcare workflows for t

BIM Information Extraction Through LLM-based Adaptive Exploration

Model ReleasesDGX agent

arXiv:2605.01698v1 Announce Type: new Abstract: BIM models provide structured representations of building geometry, semantics, and topology, yet extracting specific information from them remains remar

Combining Trained Models in Reinforcement Learning

SafetyDGX agent

arXiv:2605.02159v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) has delivered strong results in domains such as Atari and Go, but it still suffers from high sample cost and weak tran

Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation

Model ReleasesDGX agent

arXiv:2605.01448v1 Announce Type: cross Abstract: Cross-task generalization is a core challenge in open-world robotic manipulation, and the key lies in extracting transferable manipulation knowledge f

DeepStage: Learning Autonomous Defense Policies Against Multi-Stage APT Campaigns

SafetyDGX agent

arXiv:2603.16969v2 Announce Type: replace-cross Abstract: This paper presents DeepStage, a deep reinforcement learning (DRL) framework for adaptive and stage-aware defense against Advanced Persistent

DynoSLAM: Dynamic SLAM with Generative Graph Neural Networks for Real-World Social Navigation

Local AiDGX agent

arXiv:2605.02759v1 Announce Type: cross Abstract: Traditional Simultaneous Localization and Mapping (SLAM) algorithms rely heavily on the static environment assumption, which severely limits their app

Experience Constrained Hierarchical Federated Reinforcement Learning for Large-scale UAV Teams in Hazardous Environments

SafetyDGX agent

arXiv:2605.02165v1 Announce Type: new Abstract: Conventional federated learning assumes that greater learner participation improves training performance, by leveraging abundant, independently generate

G-reasoner: Foundation Models for Unified Reasoning over Graph-structured Knowledge

Model ReleasesDGX agent

arXiv:2509.24276v4 Announce Type: replace Abstract: Large language models (LLMs) excel at complex reasoning but remain limited by static and incomplete parametric knowledge. Retrieval-augmented genera

Green Energy Management for Sustainable Data Centers Using Deep Reinforcement Learning

SafetyDGX agent

arXiv:2507.21153v2 Announce Type: replace Abstract: The exponential growth of digital services has positioned data centers among the most energy-intensive infrastructures in the modern economy, raisin

High entropy leads to symmetry equivariant policies in Dec-POMDPs

SafetyDGX agent

arXiv:2511.22581v3 Announce Type: replace Abstract: We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization

Hybrid Quantum Reinforcement Learning with QAOA for Improved Vehicle Routing Optimization

SafetyDGX agent

arXiv:2605.01574v1 Announce Type: new Abstract: Vehicle Routing Problem (VRP) is one of the most complex NP-hard combinatorial optimization problem in transportation and logistics that requires a dyna

i have yet to meet a single person who feels like claude code is getting exponentially better on some kind of fast take off

Model ReleasesDGX agent

i have yet to meet a single person who feels like claude code is getting exponentially better on some kind of fast take off Anthropic pays $750K/ year per senior engineer. The creator of Claude Code j

Knowledge-Based Design Requirements for Generative Social Robots in Higher Education

SafetyDGX agent

arXiv:2602.12873v4 Announce Type: replace-cross Abstract: Generative social robots (GSRs) powered by large language models enable adaptive, conversational tutoring but also introduce risks such as mis

MolViBench: Evaluating LLMs on Molecular Vibe Coding

Model ReleasesDGX agent

arXiv:2605.02351v1 Announce Type: new Abstract: Molecular Vibe Coding, a paradigm where chemists interact with LLMs to generate executable programs for molecular tasks, has emerged as a flexible alter

OpenAI GPT-5 System Card

Model ReleasesDGX agent

arXiv:2601.03267v2 Announce Type: replace Abstract: This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers

Orchestrating Spatial Semantics via a Zone-Graph Paradigm for Intricate Indoor Scene Generation

Model ReleasesDGX agent

arXiv:2605.02537v1 Announce Type: new Abstract: Autonomous 3D indoor scene synthesis breaks down in non-convex rooms with tightly coupled spatial constraints. Data-driven generators lack topological p

Our AI started a cafe in Stockholm

ApplicationsDGX agent

Our AI started a cafe in Stockholm Andon Labs previously started an AI-run retail store in San Francisco. Now they're running a similar experiment in Stockholm, Sweden, only this time it's a cafe. The

PACE: Parameter Change for Unsupervised Environment Design

Model ReleasesDGX agent

arXiv:2605.01358v1 Announce Type: new Abstract: Unsupervised Environment Design (UED) offers a promising paradigm for improving reinforcement learning generalization by adaptively shaping training env

RAST-MoE-RL: A Regime-Aware Spatio-Temporal MoE Framework for Deep Reinforcement Learning in Ride-Hailing

ApplicationsDGX agent

arXiv:2512.13727v2 Announce Type: replace Abstract: Ride-hailing platforms face the challenge of balancing passenger waiting times with overall system efficiency under highly uncertain supply-demand c

Reliability-Oriented Multilingual Orthopedic Diagnosis: A Domain-Adaptive Modeling and a Conceptual Validation Framework

SafetyDGX agent

arXiv:2605.02266v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly proposed for clinical decision support including multilingual diagnosis in low-resource settings. However,

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving

HardwareDGX agent

arXiv:2605.01708v1 Announce Type: cross Abstract: Contemporary systems serving large language models (LLMs) have adopted prefill-decode disaggregation to better load-balance between the compute-bound

STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

SafetyDGX agent

arXiv:2605.02122v1 Announce Type: new Abstract: Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fr

STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Storie

Model ReleasesDGX agent

arXiv:2601.08510v3 Announce Type: replace Abstract: Movie screenplays are rich long-form narratives that interleave complex character relationships, temporally ordered events, and dialogue-driven inte

SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation

Model ReleasesDGX agent

arXiv:2508.15658v5 Announce Type: replace Abstract: The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show pr

toodles from mickey mouse clubhouse was weirdly ahead of its time wake phrase: mickey mouse clubhouse launched in 2006, and “oh toodles” tra…

Model ReleasesDGX agent

toodles from mickey mouse clubhouse was weirdly ahead of its time wake phrase: mickey mouse clubhouse launched in 2006, and “oh toodles” trained toddlers on the assistant wake phrase years before siri

4 May 2026

Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments

ResearchDGX agent

arXiv:2505.09901v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate or automate human behavior in complex sequential decision-making settings. A na

Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations

SafetyDGX agent

arXiv:2512.20260v5 Announce Type: replace Abstract: Weakly-Supervised Camouflaged Object Detection (WSCOD) aims to locate and segment objects that are visually concealed within their surrounding scene

Decentralized Proximal Stochastic Gradient Langevin Dynamics

SafetyDGX agent

arXiv:2605.00723v1 Announce Type: cross Abstract: We propose Decentralized Proximal Stochastic Gradient Langevin Dynamics (DE-PSGLD), a decentralized Markov chain Monte Carlo (MCMC) algorithm for samp

FollowTable: A Benchmark for Instruction-Following Table Retrieval

Model ReleasesDGX agent

arXiv:2605.00400v1 Announce Type: cross Abstract: Table Retrieval (TR) has traditionally been formulated as an ad-hoc retrieval problem, where relevance is primarily determined by topical semantic sim

Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding

SafetyDGX agent

arXiv:2605.00642v1 Announce Type: cross Abstract: Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capabili

Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values

SafetyDGX agent

arXiv:2605.00762v1 Announce Type: new Abstract: We propose a new framework for meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback (BCMAB-FBF). Unlike semi-ba

Minimizing Human Intervention in Online Classification

Model ReleasesDGX agent

arXiv:2510.23557v2 Announce Type: replace-cross Abstract: Training or fine-tuning large language model (LLM)-based systems often requires costly human feedback, yet there is limited understanding of h

Optimizing Resource-Constrained Non-Pharmaceutical Interventions for Multi-Cluster Outbreak Control Using Hierarchical Reinforcement Learning

SafetyDGX agent

arXiv:2603.19397v2 Announce Type: replace Abstract: Non-pharmaceutical interventions (NPIs), such as diagnostic testing and quarantine, are crucial for controlling infectious disease outbreaks but are

Redis Array Playground

Model ReleasesDGX agent

Tool: Redis Array Playground Salvatore Sanfilippo submitted a PR adding a new data type - arrays - to Redis. The new commands are ARCOUNT, ARDEL, ARDELRANGE, ARGET, ARGETRANGE, ARGREP, ARINFO, ARINSER

Reinforcement Learning with LLM-Guided Action Spaces for Synthesizable Lead Optimization

SafetyDGX agent

arXiv:2604.07669v2 Announce Type: replace Abstract: Lead optimization in drug discovery requires improving therapeutic properties while ensuring that molecular modifications correspond to feasible syn

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

SafetyDGX agent

arXiv:2605.00380v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diver

Semantic Level of Detail for Knowledge Graphs: Discovering Abstraction Boundaries via Spectral Heat Diffusion

Model ReleasesDGX agent

arXiv:2603.08965v2 Announce Type: replace Abstract: Graph-structured knowledge systems -- from knowledge graphs to GraphRAG pipelines -- organize information into hierarchical communities, yet lack a

This is why I keep going on about LangGraph 1.2 alpha is looking great: node-level error handlers, timeouts, improved performance, and stron…

ApplicationsDGX agent

This is why I keep going on about LangGraph 1.2 alpha is looking great: node-level error handlers, timeouts, improved performance, and stronger typing LangGraph is the runtime layer you want underneat

this one is doing v well btw if you want the popular vote filter on the firehose of all the things @patrickdebois was one of the track keyno…

ToolsDGX agent

this one is doing v well btw if you want the popular vote filter on the firehose of all the things @patrickdebois was one of the track keynotes i gave a 'blank check' to based on his sincere support s

World Model for Robot Learning: A Comprehensive Survey

SafetyDGX agent

arXiv:2605.00080v1 Announce Type: cross Abstract: World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They s

2 May 2026

Are AIs about to subjugate humanity? In the debate about catastrophic AI risk, evolutionary scenarios receive far too little attention compa…

SafetyDGX agent

Are AIs about to subjugate humanity? In the debate about catastrophic AI risk, evolutionary scenarios receive far too little attention compared to largely speculative arguments about “instrumental con

Grok Voice is used by Starlink right now

Model ReleasesDGX agent

Grok Voice is used by Starlink right now Grok Voice brutally dominates the top of the τ-voice Bench Grok scores 67.3%, while Gemini sits at 43.8% and GPT Realtime at 35.3% This is a massive lead over

I was quoted a couple times in this Atlantic article, but that isn’t (the only) reason I think it is good. It lays out the reasons why we wh…

ApplicationsDGX agent

I was quoted a couple times in this Atlantic article, but that isn’t (the only) reason I think it is good. It lays out the reasons why we whipsawed from “AI is a bubble” to “there are not enough data

People are really enjoying our full workshops showing end to end walkthroughs of real production workflows! This is a rare double header wit…

SafetyDGX agent

People are really enjoying our full workshops showing end to end walkthroughs of real production workflows! This is a rare double header with @braintrust's Giran Moodley and @OussamaHaff walking thoug

1 May 2026

Apple may take 'several months' to catch up to Mac mini and Studio demand

IndustryDGX agent

During its Q2 2026 earnings call, Apple CEO Tim Cook stated that Mac mini and Mac Studio may take several months to reach supply-demand balance. Apple underestimated demand for these products, which h

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR

Model ReleasesDGX agent

arXiv:2604.27543v1 Announce Type: new Abstract: Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into sh

ChipLingo: A Systematic Training Framework for Large Language Models in EDA

Model ReleasesDGX agent

arXiv:2604.27415v1 Announce Type: new Abstract: With the rapid advancement of semiconductor technology, Electronic Design Automation (EDA) has become an increasingly knowledge-intensive and document-d

Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations

SafetyDGX agent

arXiv:2604.27372v1 Announce Type: cross Abstract: This paper investigates the continuous-time counterpart of the Q-function for entropy-regularized mean-field control (MFC) with controlled common nois

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning…

SafetyDGX agent

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning fixes get bolted on at post-training. By then, the patterns

From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction

Model ReleasesDGX agent

arXiv:2604.27906v1 Announce Type: new Abstract: Persistent AI memory is often reduced to a retrieval problem: store prior interactions as text, embed them, and ask the model to recover relevant contex

IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development

ApplicationsDGX agent

arXiv:2604.16399v2 Announce Type: replace-cross Abstract: The widespread adoption of AI-assisted development tools in 2025 -- and the emergence of vibe coding, a practice of generating complete applic

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions

Model ReleasesDGX agent

arXiv:2604.27763v1 Announce Type: new Abstract: The emergence of Large Language Models (LLMs) offers a transformative interface for Web3, yet existing benchmarks fail to capture the complexity of tran

KellyBench: A Benchmark for Long-Horizon Sequential Decision Making

Model ReleasesDGX agent

arXiv:2604.27865v1 Announce Type: new Abstract: Language models are saturating benchmarks for procedural tasks with narrow objectives. But they are increasingly being deployed in long-horizon, non-sta

Knowledge Graph Representations for LLM-Based Policy Compliance Reasoning

SafetyDGX agent

arXiv:2604.27713v1 Announce Type: new Abstract: The risks posed by AI features are increasing as they are rapidly integrated into software applications. In response, regulations and standards for safe

Language Models Refine Mechanical Linkage Designs Through Symbolic Reflection and Modular Optimisation

Model ReleasesDGX agent

arXiv:2604.27962v1 Announce Type: new Abstract: Designing mechanical linkages involves combinatorial topology selection and continuous parameter fitting. We show that language models can systematicall

← Previous
1…281282283284285…294
Next →