AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,005 results
12 May 2026

Ace-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution

AgentsDGX agent

arXiv:2605.08887v1 Announce Type: new Abstract: Self-evolving agents present a promising path toward continual adaptation by distilling task interactions into reusable knowledge artifacts. In practice

Agent-Sentry: Bounding LLM Agents via Execution Provenance

SafetyDGX agent

arXiv:2603.22868v2 Announce Type: replace-cross Abstract: Agentic computing systems, while immensely capable, raise serious security, privacy, and safety concerns. A key issue is that the full set of

Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

AgentsDGX agent

arXiv:2605.08956v1 Announce Type: new Abstract: A growing body of work pursues AI scientists capable of end-to-end autonomous scientific discovery. This position paper argues that although they alread

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Alignment as Jurisprudence

SafetyDGX agent

arXiv:2605.08416v1 Announce Type: new Abstract: Jurisprudence, the study of how judges should properly decide cases, and alignment, the science of getting AI models to conform to human values, share a

An agentic framework for gravitational-wave counterpart association in the multi-messenger era

AgentsDGX agent

arXiv:2605.10584v1 Announce Type: cross Abstract: With the detection of gravitational waves (GWs), multi-messenger astronomy has opened a new window for advancing our understanding of astrophysics, de

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs

ResearchDGX agent

arXiv:2509.08031v3 Announce Type: replace-cross Abstract: Large Audio Language Models (LALMs) are rapidly advancing, but evaluating them remains challenging due to inefficient and non-standardized too

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards

SafetyDGX agent

arXiv:2511.14045v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a core training stage in recent large language models (LLMs). Its reliance on

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

Model ReleasesDGX agent

arXiv:2605.08761v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles,

Causal Parametric Drift Simulation: A Digital Twin Framework for Classifier Robustness Evaluation

ResearchDGX agent

arXiv:2605.09663v1 Announce Type: cross Abstract: Machine learning classifiers in dynamic environments face concept drift -- changes in the data-generating process that degrade performance. Convention

Conformity Generates Collective Misalignment in AI Agents Societies

SafetyDGX agent

arXiv:2605.10721v1 Announce Type: cross Abstract: Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability

AgentsDGX agent

arXiv:2605.10516v1 Announce Type: new Abstract: This paper establishes a rigorous measurement science for AI agent reliability, providing a foundational framework for quantifying consistency under sem

CONTRA: Conformal Prediction Region via Normalizing Flow Transformation

ResearchDGX agent

arXiv:2605.08561v1 Announce Type: cross Abstract: Density estimation and reliable prediction regions for outputs are crucial in supervised and unsupervised learning. While conformal prediction effecti

Counterfactual Stress Testing for Image Classification Models

ApplicationsDGX agent

arXiv:2605.10894v1 Announce Type: new Abstract: Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner hardwa

Cplus2ASP: Computing Action Language C+ in Answer Set Programming

ResearchDGX agent

arXiv:2605.09528v1 Announce Type: new Abstract: We present Version 2 of system Cplus2ASP, which implements the definite fragment of action language C+. Its input language is fully compatible with the

DataArc-SynData-Toolkit: A Unified Closed-Loop Framework for Multi-Path, Multimodal, and Multilingual Data Synthesis

ApplicationsDGX agent

arXiv:2605.08138v1 Announce Type: new Abstract: Synthetic data has emerged as a crucial solution to the data scarcity bottleneck in large language models (LLMs), particularly for specialized domains a

Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents

ResearchDGX agent

arXiv:2605.08442v1 Announce Type: cross Abstract: Persistent memory attacks against LLM agents achieve high attack success rates against open-source models. In these attacks, malicious instructions in

Delivering Science as a Service: Sci-Orchestra's Cloud-Native Approach to HPC

AgentsDGX agent

arXiv:2605.08396v1 Announce Type: new Abstract: The increasing complexity of modern computational environments often burdens researchers with infrastructure management, authentication protocols, and c

Discriminative Span as a Predictor of Synthetic Data Utility via Classifier Reconstruction

ApplicationsDGX agent

arXiv:2605.09697v1 Announce Type: new Abstract: In many real-world computer vision applications, including medical imaging and industrial inspection, binary classification tasks are characterized by a

Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas

TutorialsDGX agent

arXiv:2605.09159v1 Announce Type: new Abstract: Recent work shows that large language models (LLMs) encode behavioural traits ('personas') as linear directions in activation space, often called 'perso

Do you actually own your document parsing infrastructure? 👀 At @llama_index, we wanted to make that easier, so we built 𝗹𝗶𝘁𝗲𝗽𝗮𝗿𝘀𝗲-…

Model ReleasesDGX agent

Do you actually own your document parsing infrastructure? 👀 At @llama_index, we wanted to make that easier, so we built 𝗹𝗶𝘁𝗲𝗽𝗮𝗿𝘀𝗲-𝘀𝗲𝗿𝘃𝗲𝗿, a lightweight HTTP backend built on top of LiteParse that can

Efficient Statistics With Unknown Truncation, Polynomial Time Algorithms, Beyond Gaussians

ResearchDGX agent

arXiv:2410.01656v2 Announce Type: replace-cross Abstract: We study the estimation of distributional parameters when samples are shown only if they fall in some unknown set S subseteq R^d. Kontonis, Tz

Enabling Structure-Only Initialization and Out-of-Distribution Generalization in GNN-based Molecular Dynamics Simulators

ResearchDGX agent

arXiv:2605.09495v1 Announce Type: cross Abstract: Machine learning-based simulators offer the potential to model the dynamics of complex systems more efficiently than classical approaches, while retai

Engineering Robustness into Personal Agents with the AI Workflow Store

AgentsDGX agent

arXiv:2605.10907v1 Announce Type: cross Abstract: The dominant paradigm for AI agents is an 'on-the-fly' loop in which agents synthesize plans and execute actions within seconds or minutes in response

ERIS: Enhancing Privacy and Scalability in Federated Learning via Federated Shard Aggregation

Model ReleasesDGX agent

arXiv:2602.08617v2 Announce Type: replace Abstract: Scaling Federated Learning (FL) to billion-parameter models forces a challenging trade-off between privacy, scalability, and model utility. Existing

Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?

Model ReleasesDGX agent

arXiv:2603.00166v2 Announce Type: replace-cross Abstract: Recent advances in generative AI have shown human-level performance in complex content creation. However, we identify a 'Paradox of Simplicity

Extended Wasserstein-GAN Approach to Causal Distribution Learning: Density-Free Estimation and Minimax Optimality

SafetyDGX agent

arXiv:2605.10206v1 Announce Type: cross Abstract: Distributional causal inference requires estimating not only average treatment effects but also interventional outcome distributions, including quanti

Factual recall in linear associative memories: sharp asymptotics and mechanistic insights

ResearchDGX agent

arXiv:2605.10795v1 Announce Type: cross Abstract: Large language models demonstrate remarkable ability in factual recall, yet the fundamental limits of storing and retrieving input--output association

FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning

AgentsDGX agent

arXiv:2605.09932v1 Announce Type: new Abstract: Large language models can now process increasingly long inputs, yet their ability to effectively use information spread across long contexts remains lim

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

ApplicationsDGX agent

arXiv:2605.10834v1 Announce Type: new Abstract: AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perfor

From pre-training to downstream performance: Does domain-specific pre-training make sense?

ResearchDGX agent

arXiv:2605.08819v1 Announce Type: new Abstract: Deep learning techniques have revolutionised medical imaging, improving diagnostic accuracy and enabling both more accurate and earlier disease detectio

General Agent Evaluation

Model ReleasesDGX agent

arXiv:2602.22953v2 Announce Type: replace Abstract: General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measur

GPU-Accelerated Synthesis of Mixed-Boolean Arithmetic: Beyond Caching

HardwareDGX agent

arXiv:2605.08243v1 Announce Type: cross Abstract: Synthesizing Mixed-Boolean Arithmetic (MBA) expressions from input-output examples is central to program deobfuscation and also useful for compiler op

Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models

Model ReleasesDGX agent

arXiv:2603.16253v2 Announce Type: replace-cross Abstract: Vision-language process reward models (VL-PRMs) are increasingly used to score intermediate reasoning steps and rerank candidates under test-t

Harmonized Feature Conditioning and Frequency-Prompt Personalization for Multi-Rater Medical Segmentation

ResearchDGX agent

arXiv:2605.08210v1 Announce Type: new Abstract: Multi-rater medical image segmentation captures the inherent ambiguity of clinical interpretation, where diagnostic boundaries vary across experts and i

Hierarchical Prompting with Dual LLM Modules for Robotic Task and Motion Planning

AgentsDGX agent

arXiv:2605.08330v1 Announce Type: new Abstract: We present a hierarchical language-driven framework for robotic task and motion planning to improve natural, intuitive human-robot interaction in servic

How Imgix processes 8 billion images daily with G4 VMs powered by NVIDIA Blackwell

HardwareDGX agent

The modern web is extremely visual. People are busy and easily-distracted, and smart companies know they have just seconds to attract would-be customers with compelling images, videos, animations, and

How Sapu Indexed 28 Million PubMed Abstracts to Accelerate Cancer Research with Qdrant

ApplicationsDGX agent

Sapu is an early-stage biopharmaceutical company developing treatments for hard-to-treat cancers. From its San Diego facility, the team is pioneering a nanomedicine pipeline that takes existing FDA-ap

'It is categorically *not* a victory for pure LLMs; it’s a victory for borrowing from classical AI and CS to move *beyond* pure LLMs.' ☝️ Co…

Model ReleasesDGX agent

'It is categorically *not* a victory for pure LLMs; it’s a victory for borrowing from classical AI and CS to move *beyond* pure LLMs.' ☝️ Compound AI Systems alllll the way downnn 🤩🤯🤩 Claude Code (sti

Kintsugi: Learning Policies by Repairing Executable Knowledge Bases

Local AiDGX agent

arXiv:2605.09487v1 Announce Type: new Abstract: Modern embodied agents achieve impressive performance, but their task knowledge is often stored in neural weights, latent state, or prompt-bound memory,

LAQuant: A Simple Overhead-free Large Reasoning Model Quantization by Layer-wise Lookahead Loss

SafetyDGX agent

arXiv:2605.08755v1 Announce Type: new Abstract: Large reasoning models (LRMs) reach competition-level math and coding accuracy via long autoregressive decoding, making per-token decoding cost a primar

llm 0.32a2

Model ReleasesDGX agent

Release: llm 0.32a2 A bunch of useful stuff in this LLM alpha, but the most important detail is this one: Most reasoning-capable OpenAI models now use the /v1/responses endpoint instead of /v1/chat/co

Machine Learning Research Has Outpaced Its Communication Norms and NeurIPS Should Act

ResearchDGX agent

arXiv:2605.08889v1 Announce Type: cross Abstract: Machine learning research has grown exponentially while its communication norms have not. We argue NeurIPS should adopt explicit, measurable writing s

MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service

SafetyDGX agent

arXiv:2605.08527v1 Announce Type: cross Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has significantly improved the reasoning capabilities of large language models (LLMs), particula

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

Model ReleasesDGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

Measuring and Decomposing Mode Separation via the Canonical Diffusion

ResearchDGX agent

arXiv:2605.08777v1 Announce Type: cross Abstract: Mode separation, namely how sharply a distribution fragments into barrier-separated clusters, is a fundamental geometric property of densities, diffic

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents

Model ReleasesDGX agent

arXiv:2605.09863v1 Announce Type: cross Abstract: Production LLM coding agents drift over long sessions: they forget user-specified constraints, slip into mistakes the user already flagged, and confab

Need document parsing that stays fully local and private? 👀 Meet liteparse-server, a self-hostable, open-source HTTP server for parsing doc…

Model ReleasesDGX agent

Need document parsing that stays fully local and private? 👀 Meet liteparse-server, a self-hostable, open-source HTTP server for parsing documents and generating screenshots from PDFs, Office files, an

NeuroGAN-3D: Enhancing Intrinsic Functional Brain Networks via High-Fidelity 3D Generative Super-Resolution

Local AiDGX agent

arXiv:2605.08373v1 Announce Type: cross Abstract: Recent advances in neuroimaging have deepened our understanding of the brain's complex functional and structural organization. Among these, functional

On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective

Local AiDGX agent

arXiv:2605.08368v1 Announce Type: new Abstract: Debates about large language model post-training often treat supervised fine-tuning (SFT) as imitation and reinforcement learning (RL) as discovery. But

PAAC: Privacy-Aware Agentic Device-Cloud Collaboration

Local AiDGX agent

arXiv:2605.08646v1 Announce Type: cross Abstract: Large language model (LLM) agents face a structural tension: cloud agents provide strong reasoning but expose user data, while on-device agents preser

PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents

Model ReleasesDGX agent

arXiv:2605.10341v1 Announce Type: new Abstract: A LaTeX manuscript that compiles without error is not necessarily publication-ready. The resulting PDFs frequently suffer from misplaced floats, overflo

Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes

Model ReleasesDGX agent

arXiv:2602.00593v2 Announce Type: replace Abstract: Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and e

Position: Life-Logging Video Streams Make the Privacy-Utility Trade-off Inevitable

ResearchDGX agent

arXiv:2605.10404v1 Announce Type: new Abstract: With the growing prevalence of always-on hardware such as smart glasses, body cameras, and home security systems, life-logging visual sensing is becomin

PRISM: Fast Online LLM Serving via Scheduling-Memory Co-design

AgentsDGX agent

arXiv:2605.08581v1 Announce Type: new Abstract: Modern online large language model (LLM) services, such as Retrieval-Augmented Generation (RAG) and agent systems, increasingly expose two prominent cha

Privacy Auditing Synthetic Data Release through Local Likelihood Attacks

Model ReleasesDGX agent

arXiv:2508.21146v2 Announce Type: replace Abstract: Auditing the privacy leakage of synthetic data is an important but unresolved problem. Existing privacy auditing frameworks for synthetic data rely

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic

ResearchDGX agent

arXiv:2605.09157v1 Announce Type: cross Abstract: Mixture policies theoretically offer greater flexibility than unimodal policies in continuous action reinforcement learning, but the practical benefit

RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

SafetyDGX agent

arXiv:2605.10899v1 Announce Type: new Abstract: Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyo

SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

SafetyDGX agent

arXiv:2605.08334v1 Announce Type: new Abstract: We present SalesSim, a framework and testbed for evaluating the ability of Multimodal Large Language Models (MLLMs) to simulate realistic, persona-drive

SAR-RAG: ATR Visual Question Answering by Semantic Search, Retrieval, and MLLM Generation

AgentsDGX agent

arXiv:2602.04712v2 Announce Type: replace-cross Abstract: We present a visual-context image-retrieval-augmented generation (ImageRAG)- assisted AI agent for automatic target recognition (ATR) of synth

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning

AgentsDGX agent

arXiv:2605.09423v1 Announce Type: new Abstract: LLM/VLM-based digital agents have advanced rapidly thanks to scalable sandboxes for coding, web navigation, and computer use, which provide rich interac

← Previous
1…143144145146147…167
Next →