AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,005 results
Agents

Ace-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution

DGX agent

arXiv:2605.08887v1 Announce Type: new Abstract: Self-evolving agents present a promising path toward continual adaptation by distilling task interactions into reusable knowledge artifacts. In practice

agentsarxiv-cs-ai
12 May 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Agent-Sentry: Bounding LLM Agents via Execution Provenance

DGX agent

arXiv:2603.22868v2 Announce Type: replace-cross Abstract: Agentic computing systems, while immensely capable, raise serious security, privacy, and safety concerns. A key issue is that the full set of

safetyarxiv-cs-ai
12 May 2026
Agents

Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

DGX agent

arXiv:2605.08956v1 Announce Type: new Abstract: A growing body of work pursues AI scientists capable of end-to-end autonomous scientific discovery. This position paper argues that although they alread

agentsarxiv-cs-ai
12 May 2026
Safety

Alignment as Jurisprudence

DGX agent

arXiv:2605.08416v1 Announce Type: new Abstract: Jurisprudence, the study of how judges should properly decide cases, and alignment, the science of getting AI models to conform to human values, share a

safetyarxiv-cs-ai
12 May 2026
Agents

An agentic framework for gravitational-wave counterpart association in the multi-messenger era

DGX agent

arXiv:2605.10584v1 Announce Type: cross Abstract: With the detection of gravitational waves (GWs), multi-messenger astronomy has opened a new window for advancing our understanding of astrophysics, de

agentsarxiv-cs-ai
12 May 2026
Research

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs

DGX agent

arXiv:2509.08031v3 Announce Type: replace-cross Abstract: Large Audio Language Models (LALMs) are rapidly advancing, but evaluating them remains challenging due to inefficient and non-standardized too

researcharxiv-cs-ai
12 May 2026
Safety

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards

DGX agent

arXiv:2511.14045v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a core training stage in recent large language models (LLMs). Its reliance on

safetyarxiv-cs-ai
12 May 2026
Model Releases

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

DGX agent

arXiv:2605.08761v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles,

model-releasesarxiv-cs-lg
12 May 2026
Research

Causal Parametric Drift Simulation: A Digital Twin Framework for Classifier Robustness Evaluation

DGX agent

arXiv:2605.09663v1 Announce Type: cross Abstract: Machine learning classifiers in dynamic environments face concept drift -- changes in the data-generating process that degrade performance. Convention

researcharxiv-cs-ai
12 May 2026
Safety

Conformity Generates Collective Misalignment in AI Agents Societies

DGX agent

arXiv:2605.10721v1 Announce Type: cross Abstract: Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate

safetyarxiv-cs-cl
12 May 2026
Agents

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability

DGX agent

arXiv:2605.10516v1 Announce Type: new Abstract: This paper establishes a rigorous measurement science for AI agent reliability, providing a foundational framework for quantifying consistency under sem

agentsarxiv-cs-ai
12 May 2026
Research

CONTRA: Conformal Prediction Region via Normalizing Flow Transformation

DGX agent

arXiv:2605.08561v1 Announce Type: cross Abstract: Density estimation and reliable prediction regions for outputs are crucial in supervised and unsupervised learning. While conformal prediction effecti

researcharxiv-cs-lg
12 May 2026
Applications

Counterfactual Stress Testing for Image Classification Models

DGX agent

arXiv:2605.10894v1 Announce Type: new Abstract: Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner hardwa

applicationsarxiv-cs-cv
12 May 2026
Research

Cplus2ASP: Computing Action Language C+ in Answer Set Programming

DGX agent

arXiv:2605.09528v1 Announce Type: new Abstract: We present Version 2 of system Cplus2ASP, which implements the definite fragment of action language C+. Its input language is fully compatible with the

researcharxiv-cs-ai
12 May 2026
Applications

DataArc-SynData-Toolkit: A Unified Closed-Loop Framework for Multi-Path, Multimodal, and Multilingual Data Synthesis

DGX agent

arXiv:2605.08138v1 Announce Type: new Abstract: Synthetic data has emerged as a crucial solution to the data scarcity bottleneck in large language models (LLMs), particularly for specialized domains a

applicationsarxiv-cs-lg
12 May 2026
Research

Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents

DGX agent

arXiv:2605.08442v1 Announce Type: cross Abstract: Persistent memory attacks against LLM agents achieve high attack success rates against open-source models. In these attacks, malicious instructions in

researcharxiv-cs-ai
12 May 2026
Agents

Delivering Science as a Service: Sci-Orchestra's Cloud-Native Approach to HPC

DGX agent

arXiv:2605.08396v1 Announce Type: new Abstract: The increasing complexity of modern computational environments often burdens researchers with infrastructure management, authentication protocols, and c

agentsarxiv-cs-cv
12 May 2026
Applications

Discriminative Span as a Predictor of Synthetic Data Utility via Classifier Reconstruction

DGX agent

arXiv:2605.09697v1 Announce Type: new Abstract: In many real-world computer vision applications, including medical imaging and industrial inspection, binary classification tasks are characterized by a

applicationsarxiv-cs-cv
12 May 2026
Tutorials

Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas

DGX agent

arXiv:2605.09159v1 Announce Type: new Abstract: Recent work shows that large language models (LLMs) encode behavioural traits ('personas') as linear directions in activation space, often called 'perso

tutorialsarxiv-cs-ai
12 May 2026
Model Releases

Do you actually own your document parsing infrastructure? 👀 At @llama_index, we wanted to make that easier, so we built 𝗹𝗶𝘁𝗲𝗽𝗮𝗿𝘀𝗲-…

DGX agent

Do you actually own your document parsing infrastructure? 👀 At @llama_index, we wanted to make that easier, so we built 𝗹𝗶𝘁𝗲𝗽𝗮𝗿𝘀𝗲-𝘀𝗲𝗿𝘃𝗲𝗿, a lightweight HTTP backend built on top of LiteParse that can

model-releasesjerry-liu--x
12 May 2026
Research

Efficient Statistics With Unknown Truncation, Polynomial Time Algorithms, Beyond Gaussians

DGX agent

arXiv:2410.01656v2 Announce Type: replace-cross Abstract: We study the estimation of distributional parameters when samples are shown only if they fall in some unknown set S subseteq R^d. Kontonis, Tz

researcharxiv-cs-lg
12 May 2026
Research

Enabling Structure-Only Initialization and Out-of-Distribution Generalization in GNN-based Molecular Dynamics Simulators

DGX agent

arXiv:2605.09495v1 Announce Type: cross Abstract: Machine learning-based simulators offer the potential to model the dynamics of complex systems more efficiently than classical approaches, while retai

researcharxiv-cs-lg
12 May 2026
Agents

Engineering Robustness into Personal Agents with the AI Workflow Store

DGX agent

arXiv:2605.10907v1 Announce Type: cross Abstract: The dominant paradigm for AI agents is an 'on-the-fly' loop in which agents synthesize plans and execute actions within seconds or minutes in response

agentsarxiv-cs-ai
12 May 2026
Model Releases

ERIS: Enhancing Privacy and Scalability in Federated Learning via Federated Shard Aggregation

DGX agent

arXiv:2602.08617v2 Announce Type: replace Abstract: Scaling Federated Learning (FL) to billion-parameter models forces a challenging trade-off between privacy, scalability, and model utility. Existing

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?

DGX agent

arXiv:2603.00166v2 Announce Type: replace-cross Abstract: Recent advances in generative AI have shown human-level performance in complex content creation. However, we identify a 'Paradox of Simplicity

model-releasesarxiv-cs-ai
12 May 2026
Safety

Extended Wasserstein-GAN Approach to Causal Distribution Learning: Density-Free Estimation and Minimax Optimality

DGX agent

arXiv:2605.10206v1 Announce Type: cross Abstract: Distributional causal inference requires estimating not only average treatment effects but also interventional outcome distributions, including quanti

safetyarxiv-cs-lg
12 May 2026
Research

Factual recall in linear associative memories: sharp asymptotics and mechanistic insights

DGX agent

arXiv:2605.10795v1 Announce Type: cross Abstract: Large language models demonstrate remarkable ability in factual recall, yet the fundamental limits of storing and retrieving input--output association

researcharxiv-cs-lg
12 May 2026
Agents

FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning

DGX agent

arXiv:2605.09932v1 Announce Type: new Abstract: Large language models can now process increasingly long inputs, yet their ability to effectively use information spread across long contexts remains lim

agentsarxiv-cs-cl
12 May 2026
Applications

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

DGX agent

arXiv:2605.10834v1 Announce Type: new Abstract: AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perfor

applicationsarxiv-cs-ai
12 May 2026
Research

From pre-training to downstream performance: Does domain-specific pre-training make sense?

DGX agent

arXiv:2605.08819v1 Announce Type: new Abstract: Deep learning techniques have revolutionised medical imaging, improving diagnostic accuracy and enabling both more accurate and earlier disease detectio

researcharxiv-cs-cv
12 May 2026
Model Releases

General Agent Evaluation

DGX agent

arXiv:2602.22953v2 Announce Type: replace Abstract: General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measur

model-releasesarxiv-cs-ai
12 May 2026
Hardware

GPU-Accelerated Synthesis of Mixed-Boolean Arithmetic: Beyond Caching

DGX agent

arXiv:2605.08243v1 Announce Type: cross Abstract: Synthesizing Mixed-Boolean Arithmetic (MBA) expressions from input-output examples is central to program deobfuscation and also useful for compiler op

hardwarearxiv-cs-lg
12 May 2026
Model Releases

Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models

DGX agent

arXiv:2603.16253v2 Announce Type: replace-cross Abstract: Vision-language process reward models (VL-PRMs) are increasingly used to score intermediate reasoning steps and rerank candidates under test-t

model-releasesarxiv-cs-ai
12 May 2026
Research

Harmonized Feature Conditioning and Frequency-Prompt Personalization for Multi-Rater Medical Segmentation

DGX agent

arXiv:2605.08210v1 Announce Type: new Abstract: Multi-rater medical image segmentation captures the inherent ambiguity of clinical interpretation, where diagnostic boundaries vary across experts and i

researcharxiv-cs-cv
12 May 2026
Agents

Hierarchical Prompting with Dual LLM Modules for Robotic Task and Motion Planning

DGX agent

arXiv:2605.08330v1 Announce Type: new Abstract: We present a hierarchical language-driven framework for robotic task and motion planning to improve natural, intuitive human-robot interaction in servic

agentsarxiv-cs-ro
12 May 2026
Hardware

How Imgix processes 8 billion images daily with G4 VMs powered by NVIDIA Blackwell

DGX agent

The modern web is extremely visual. People are busy and easily-distracted, and smart companies know they have just seconds to attract would-be customers with compelling images, videos, animations, and

hardwaregoogle-cloud-ai
12 May 2026
Applications

How Sapu Indexed 28 Million PubMed Abstracts to Accelerate Cancer Research with Qdrant

DGX agent

Sapu is an early-stage biopharmaceutical company developing treatments for hard-to-treat cancers. From its San Diego facility, the team is pioneering a nanomedicine pipeline that takes existing FDA-ap

applicationsqdrant
12 May 2026
Model Releases

'It is categorically *not* a victory for pure LLMs; it’s a victory for borrowing from classical AI and CS to move *beyond* pure LLMs.' ☝️ Co…

DGX agent

'It is categorically *not* a victory for pure LLMs; it’s a victory for borrowing from classical AI and CS to move *beyond* pure LLMs.' ☝️ Compound AI Systems alllll the way downnn 🤩🤯🤩 Claude Code (sti

model-releasesgary-marcus--x
12 May 2026
Local Ai

Kintsugi: Learning Policies by Repairing Executable Knowledge Bases

DGX agent

arXiv:2605.09487v1 Announce Type: new Abstract: Modern embodied agents achieve impressive performance, but their task knowledge is often stored in neural weights, latent state, or prompt-bound memory,

local-aiarxiv-cs-lg
12 May 2026
Safety

LAQuant: A Simple Overhead-free Large Reasoning Model Quantization by Layer-wise Lookahead Loss

DGX agent

arXiv:2605.08755v1 Announce Type: new Abstract: Large reasoning models (LRMs) reach competition-level math and coding accuracy via long autoregressive decoding, making per-token decoding cost a primar

safetyarxiv-cs-lg
12 May 2026
Model Releases

llm 0.32a2

DGX agent

Release: llm 0.32a2 A bunch of useful stuff in this LLM alpha, but the most important detail is this one: Most reasoning-capable OpenAI models now use the /v1/responses endpoint instead of /v1/chat/co

model-releasessimon-willison
12 May 2026
Research

Machine Learning Research Has Outpaced Its Communication Norms and NeurIPS Should Act

DGX agent

arXiv:2605.08889v1 Announce Type: cross Abstract: Machine learning research has grown exponentially while its communication norms have not. We argue NeurIPS should adopt explicit, measurable writing s

researcharxiv-cs-cl
12 May 2026
Safety

MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service

DGX agent

arXiv:2605.08527v1 Announce Type: cross Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has significantly improved the reasoning capabilities of large language models (LLMs), particula

safetyarxiv-cs-ai
12 May 2026
Model Releases

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

DGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

model-releasesarxiv-cs-ai
12 May 2026
Research

Measuring and Decomposing Mode Separation via the Canonical Diffusion

DGX agent

arXiv:2605.08777v1 Announce Type: cross Abstract: Mode separation, namely how sharply a distribution fragments into barrier-separated clusters, is a fundamental geometric property of densities, diffic

researcharxiv-cs-lg
12 May 2026
Model Releases

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents

DGX agent

arXiv:2605.09863v1 Announce Type: cross Abstract: Production LLM coding agents drift over long sessions: they forget user-specified constraints, slip into mistakes the user already flagged, and confab

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Need document parsing that stays fully local and private? 👀 Meet liteparse-server, a self-hostable, open-source HTTP server for parsing doc…

DGX agent

Need document parsing that stays fully local and private? 👀 Meet liteparse-server, a self-hostable, open-source HTTP server for parsing documents and generating screenshots from PDFs, Office files, an

model-releasesllamaindex--x
12 May 2026
Local Ai

NeuroGAN-3D: Enhancing Intrinsic Functional Brain Networks via High-Fidelity 3D Generative Super-Resolution

DGX agent

arXiv:2605.08373v1 Announce Type: cross Abstract: Recent advances in neuroimaging have deepened our understanding of the brain's complex functional and structural organization. Among these, functional

local-aiarxiv-cs-ai
12 May 2026
← Previous
1…179180181182183…209
Next →