AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,023 results
1 Jul 2026

Signed-Permutation Coordinate Transport for RMSNorm Transformers

Model ReleasesDGX agent

arXiv:2606.31963v1 Announce Type: cross Abstract: Modern LLM workflows move coordinate-indexed objects across checkpoints: steering vectors, sparse autoencoders, top-k neuron sets, attribution lists,

The best players want to be coached. You build trust, then you push hard. Now here's the thing: your agent has no skin in the game. no ego t…

AgentsDGX agent

The best players want to be coached. You build trust, then you push hard. Now here's the thing: your agent has no skin in the game. no ego to protect, no trust to earn first. So you can and should ski

Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.30801v1 Announce Type: new Abstract: Personalization algorithms determine what content users encounter on online platforms. Auditing these systems is difficult because independent auditors

When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models

Model ReleasesDGX agent

arXiv:2606.30852v1 Announce Type: new Abstract: Reasoning models spend different amounts of useful computation across instances, but it remains unclear when a learned stopping rule improves over simpl

30 Jun 2026

A causal modeling perspective on decision theory

SafetyDGX agent

arXiv:2606.29911v1 Announce Type: new Abstract: Decision theory provides a formal framework for how agents should make choices under uncertainty, drawing on ideas from philosophy, probability, and cau

A Hybrid Framework For Crypto-Ransomware Detection In Enterprise Shared Storage

Local AiDGX agent

arXiv:2606.30586v1 Announce Type: cross Abstract: Most corporate workplace environments enforce policies and technical controls that limit the storage of sensitive data on client endpoints. Consequent

A Machine-Verified Proof of a Quantum-Optimization Conjecture

Model ReleasesDGX agent

arXiv:2606.29687v1 Announce Type: cross Abstract: We report a machine-verified resolution of a problem open for over a decade in quantum optimization: the Farhi, Goldstone and Gutmann (FGG) conjecture

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

Model ReleasesDGX agent

arXiv:2606.28757v1 Announce Type: new Abstract: Generative world models hold immense promise as scalable simulators for autonomous systems, particularly for synthesizing rare but safety-critical multi

Adam's Law: Textual Frequency Law on Large Language Models

AgentsDGX agent

arXiv:2604.02176v3 Announce Type: replace Abstract: While textual frequency has been validated as relevant to human cognition in reading speed, its relatedness to Large Language Models (LLMs) is seldo

Aikido acquires Root to patch open-source software without forced upgrades

IndustryDGX agent

Belgian cybersecurity company Aikido Security NV today announced that it has acquired Root.io Inc., a company that offers patching for vulnerable open-source software at the exact versions organizatio

Anisotropy Decides Cosine vs. Rank Metrics for Text Embeddings

Model ReleasesDGX agent

arXiv:2606.29571v1 Announce Type: new Abstract: The standard way to compare two text embeddings is cosine similarity. Scattered studies report that a different metric does better, but never pin down t

Anomaly Factory 3D: A Modular Framework for Diverse Pseudo-Anomaly Synthesis in Unsupervised 3D Anomaly Detection

Local AiDGX agent

arXiv:2606.29181v1 Announce Type: cross Abstract: Detecting and localizing defects in 3D point clouds is challenging because abnormal samples are scarce and diverse, while training is often limited to

Anthropic launches Claude Sonnet 5, saying it nears Opus 4.8 performance at lower prices and is substantially better than Sonnet 4.6 for agentic work (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic launches Claude Sonnet 5, saying it nears Opus 4.8 performance at lower prices and is substantially better than Sonnet 4.6 for agentic work — Claude Sonnet 5 is built to be the mo

AutoB2G: Agentic Simulation and Reinforcement Learning for Spatio-Temporal Grid-Interactive Building Control

AgentsDGX agent

arXiv:2603.26005v2 Announce Type: replace Abstract: Grid-interactive building control has emerged as a promising approach for improving demand-side flexibility in modern power systems. Realistic studi

Bricker to BRACE: A Bracket Exposure RAW Dataset and Restoration Model for Flicker-Banding

ResearchDGX agent

arXiv:2606.29845v1 Announce Type: new Abstract: Flicker-banding (FB), arises from temporal aliasing between a camera's rolling shutter and a display's brightness modulation, degrading screen-captured

Claude Sonnet 5 is now available in Devin Desktop and Devin CLI. Sonnet 5 pairs frontier-level coding performance with a more affordable pri…

Model ReleasesDGX agent

Claude Sonnet 5 has been integrated into Devin Desktop and Devin CLI, offering advanced coding capabilities at a more competitive price point than previous models. This release represents an update to

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

Model ReleasesDGX agent

arXiv:2606.29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Ye

Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action

AgentsDGX agent

arXiv:2506.13932v3 Announce Type: replace-cross Abstract: The rise of large language models (LLMs) has led to dramatic improvements across a wide range of natural language tasks. Their performance on

COHORT: Collaborative Orchestration for Hardening via Offensive Replay on Emulated Topologies

Model ReleasesDGX agent

arXiv:2606.30479v1 Announce Type: cross Abstract: Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to that adver

Concentration bounds on response-based vector embeddings of black-box generative models

ResearchDGX agent

arXiv:2511.08307v2 Announce Type: replace-cross Abstract: Generative models, such as large language models or text-to-image diffusion models, can generate relevant responses to user-given queries. Res

Conversational Query Engine for Mixed-Modality Heterogeneous Enterprise Data Sources

Model ReleasesDGX agent

arXiv:2606.28370v1 Announce Type: cross Abstract: Enterprise business intelligence queries span structured warehouses and unstructured document repositories -- modalities with fundamentally different

Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies

SafetyDGX agent

arXiv:2606.29898v1 Announce Type: cross Abstract: Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challenges they are

Cybersecurity is the True Frontier for Generative AI Success or Failure

ResearchDGX agent

arXiv:2606.28929v1 Announce Type: cross Abstract: Cybersecurity is a real-life test-bed for many machine learning problems at once, especially when considering modern strides in using Large Language M

DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification

AgentsDGX agent

arXiv:2606.29746v1 Announce Type: new Abstract: Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge bases, remains

DeepTrans Studio: Turning Expert Interventions into Shared Team Knowledge in Agentic Translation Workflows

AgentsDGX agent

arXiv:2606.29727v1 Announce Type: new Abstract: Professional translation is often a team-based process: translators, reviewers, and project managers must coordinate terminology, legal force, and accou

DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects

Model ReleasesDGX agent

arXiv:2604.05318v2 Announce Type: replace Abstract: Harmful content detectors, particularly disinformation classifiers, are predominantly developed and evaluated on Standard American English (SAE), le

Diagnosing and Mitigating Retrieval Bottlenecks in LLM-Based Cold-Start Recommendation

Model ReleasesDGX agent

arXiv:2606.29947v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as rerankers in recommender systems, with the expectation that semantic understanding will help in

Early Estimation of Language to Latent Alignment in Diffusion Models

Model ReleasesDGX agent

arXiv:2512.08505v2 Announce Type: replace Abstract: Conditional diffusion models frequently suffer from language-image misalignments. Due to the ambiguity of intermediate noise corrupted latents, asse

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

Model ReleasesDGX agent

arXiv:2606.30219v1 Announce Type: new Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while th

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

Model ReleasesDGX agent

arXiv:2507.05257v4 Announce Type: replace-cross Abstract: Recent benchmarks for Large Language Model (LLM) agents primarily focus on evaluating reasoning, planning, and execution capabilities, while a

Factorizable Normalizing Flows for parameter-dependent density morphing

Model ReleasesDGX agent

arXiv:2606.30489v1 Announce Type: cross Abstract: Normalizing Flows excel at modeling a single fixed density, yet many problems across the sciences, such as high energy physics, instead require modeli

Fast Numbers, Slow Language: Bridging Quantitative and Qualitative Earnings Signals

ResearchDGX agent

arXiv:2606.29734v1 Announce Type: new Abstract: Earnings announcements release two types of information sequentially: quantitative surprise (numeric earnings-per-share (EPS)/revenue versus analyst est

Have your agent record video demos of its work with shot-scraper video

Model ReleasesDGX agent

shot-scraper video is a new command introduced in today's shot-scraper 1.10 release which accepts a storyboard.yml file defining a routine to run against a web application and uses Playwright to recor

Hierarchical Experimentalist Agents

Model ReleasesDGX agent

arXiv:2606.29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametr

How much of an LLM-generated clinical corpus is actually new? A production-scale measurement of content redundancy for provenance classification

Model ReleasesDGX agent

arXiv:2606.29605v1 Announce Type: new Abstract: Clinical machine learning increasingly relies on training corpora generated by large language models (LLMs) rather than annotated by clinicians, and suc

If you build with MCPs, this one is worth reading. (bookmark it) The paper covers five recurring MCP server patterns across fifteen independ…

AgentsDGX agent

If you build with MCPs, this one is worth reading. (bookmark it) The paper covers five recurring MCP server patterns across fifteen independently developed servers. That taxonomy is useful because I s

IG-Lens: Exact Additive Probability Attribution Across Transformer Layers via Telescoping Integrated Gradients

ResearchDGX agent

arXiv:2606.29693v1 Announce Type: new Abstract: We ask a simple question about decoder-only transformers: between which two layers is the probability of a predicted token actually produced? Existing l

Informational Frustration in Neural Manifolds: Shannon Bottlenecks and the Limits of Learnability

TutorialsDGX agent

arXiv:2606.30512v1 Announce Type: cross Abstract: Why overparameterised deep networks generalise so remarkably well remains one of the most stubborn open questions in machine learning theory. Classica

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

SafetyDGX agent

arXiv:2512.17504v2 Announce Type: replace-cross Abstract: Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) rema

KernelSight-LM: A Kernel-Level LLM Inference Simulator

Local AiDGX agent

arXiv:2606.28565v1 Announce Type: cross Abstract: As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, model

LAMP: Lean-based Agentic framework with MCP and Proof Repair

AgentsDGX agent

arXiv:2606.28841v1 Announce Type: cross Abstract: Large language models are increasingly capable of mathematical reasoning, but the proofs they generate are often unreliable and hard to verify. Intera

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents

Model ReleasesDGX agent

arXiv:2606.29399v1 Announce Type: new Abstract: Reviewing nuclear regulatory documents requires multi-hop reasoning across tens of thousands of pages, where judgments depend on evidence assembled acro

LLM Semantic Signaling Game and Mechanism Design: Systematic Blindness, Awareness Shaping, and Mindset Dynamics

SafetyDGX agent

arXiv:2606.29113v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate strategic interactions through natural language, making semantic control a critical element of commu

Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework

Model ReleasesDGX agent

arXiv:2606.29808v1 Announce Type: cross Abstract: Chart data extraction, which reverse-engineers data tables from chart images, is essential for reproducibility, analysis, retrieval, and redesign. Exi

Mechanistically Eliciting Latent Behaviors in Language Models

SafetyDGX agent

arXiv:2606.29604v1 Announce Type: cross Abstract: We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes. Such perturbations could help resha

Meshtryoshka: Differentiable Rendering of Real-World Scenes via Mesh Rasterization

ApplicationsDGX agent

arXiv:2606.28622v1 Announce Type: new Abstract: Differentiable rendering has emerged as a powerful approach for 3D reconstruction and novel view synthesis. State-of-the-art differentiable rendering me

MirrorCode: AI can rebuild entire programs from behavior alone

Model ReleasesDGX agent

arXiv:2606.30182v1 Announce Type: new Abstract: AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler. Ho

Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning

SafetyDGX agent

arXiv:2602.13562v2 Announce Type: replace-cross Abstract: While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety measu

ML-Powered LDAP Reconnaissance Detection using Weak Supervision

ResearchDGX agent

arXiv:2606.28917v1 Announce Type: new Abstract: Lightweight Directory Access Protocol (LDAP) is a protocol that allows users to query and modify Active Directory (AD) data. By default, all users have

Modelling Emotional Memory in Children with Tensor Networks

ApplicationsDGX agent

arXiv:2606.28470v1 Announce Type: new Abstract: We demonstrate how emotional valence influences the order-dependent structure of children's recognition memory: correct recall of a sequence of emotiona

NeuReasoner: Theory-grounded Mapping of Reasoning Elicitation Boundaries

ResearchDGX agent

arXiv:2606.29971v1 Announce Type: new Abstract: A growing body of work suggests that the reasoning capabilities of large language models are largely latent in their base form, with post-training prima

NVIDIA BioNeMo Agent Toolkit Brings Accelerated AI to Life Sciences Researchers in Claude Science

Model ReleasesDGX agent

Life sciences has entered an era of computational scale, and for more than a decade, NVIDIA has built the full GPU-accelerated computing stack — spanning hardware, frameworks, libraries, models, micro

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms

ResearchDGX agent

arXiv:2606.30625v1 Announce Type: cross Abstract: Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effectively ignori

PCP-GAN: Property-Constrained Pore-scale image reconstruction via conditional Generative Adversarial Networks

ResearchDGX agent

arXiv:2510.19465v2 Announce Type: replace Abstract: Obtaining truly representative pore-scale images that match bulk formation properties remains a fundamental challenge in subsurface characterization

Perforce launches Agentic Gateway to govern AI agents and cut token costs

Model ReleasesDGX agent

Perforce Software Inc. today expanded its Perforce Intelligence lineup with an agentic gateway for managing artificial intelligence agents, an autonomous testing platform driven by natural language an

Real-time Rendering-based Surgical Instrument Tracking via Evolutionary Optimization

ApplicationsDGX agent

arXiv:2603.11404v3 Announce Type: replace-cross Abstract: Accurate and efficient tracking of surgical instruments is fundamental for Robot-Assisted Minimally Invasive Surgery. Although vision-based ro

Recursive Self-Evolving Agents via Held-Out Selection

Model ReleasesDGX agent

arXiv:2606.28374v1 Announce Type: new Abstract: LLM agents are increasingly improved without weight updates by evolving a natural-language artifact, such as reflections, workflows, playbooks, cheatshe

ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies

SafetyDGX agent

arXiv:2606.28939v1 Announce Type: new Abstract: Behavior-cloned diffusion policies are expressive but remain vulnerable to covariate shift: small deviations from demonstrated states can compound into

Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering

Model ReleasesDGX agent

arXiv:2606.30294v1 Announce Type: new Abstract: Live product demonstrations are a recurring, high-cost activity in software organizations: a human presenter must select features, dispatch the correspo

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

AgentsDGX agent

arXiv:2606.29538v1 Announce Type: cross Abstract: Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill librari

← Previous
1…118119120121122…168
Next →