AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,620 results
Model Releases

Today we release Contrastive Neuron Attribution (CNA), a method for steering LLM behavior by identifying and ablating sparse circuits in the…

DGX agent

Today we release Contrastive Neuron Attribution (CNA), a method for steering LLM behavior by identifying and ablating sparse circuits in the MLP basis without training a sparse autoencoder, modifying

model-releasesnous-research--x
19 May 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints

DGX agent

arXiv:2602.21265v2 Announce Type: replace Abstract: We introduce ToolMATH, a math-grounded diagnostic benchmark for evaluating long-horizon tool use under controllable tool-catalog conditions. ToolMAT

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning

DGX agent

arXiv:2602.10503v2 Announce Type: replace Abstract: Pretrained on large-scale and diverse datasets, VLA models demonstrate strong generalization and adaptability as general-purpose robotic policies. H

model-releasesarxiv-cs-ro
19 May 2026
Model Releases

TPV: Parameter Perturbations Through the Lens of Test Prediction Variance

DGX agent

arXiv:2512.11089v4 Announce Type: replace-cross Abstract: We introduce test prediction variance (TPV)--the first-order sensitivity of a trained model's outputs to parameter perturbations--as a unifyin

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment

DGX agent

arXiv:2605.17342v1 Announce Type: cross Abstract: Standard RLHF relies on transitive scalar rewards, failing to capture the cyclic nature of human preferences. While some approaches like the General P

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT

DGX agent

arXiv:2605.16572v1 Announce Type: new Abstract: Automated segmentation of liver lesions on non-contrast computed tomography (NCCT) is clinically important but fundamentally challenging, particularly i

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

DGX agent

arXiv:2605.17453v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly rely on external tools to make consequential decisions, yet most existing agent-security benchmarks and defenses im

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Trustworthiness in Retrieval-Augmented Generation Systems: A Survey

DGX agent

arXiv:2409.10102v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) has quickly grown into a pivotal paradigm in the development of Large Language Models (LLMs). Although ex

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Try Anthropic Claude in Comfy today: https://links.comfy.org/4eWCvFk

DGX agent

ComfyUI announced the availability of Anthropic's Claude AI model integrated into their platform, allowing users to utilize Claude's capabilities within the ComfyUI interface. The announcement include

model-releasescomfyui--x
19 May 2026
Model Releases

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens

DGX agent

arXiv:2605.16638v1 Announce Type: new Abstract: Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning. In this paradig

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

TusoAI: Agentic Optimization for Scientific Methods

DGX agent

arXiv:2509.23986v2 Announce Type: replace Abstract: Scientific discovery is often slowed by the manual development of computational tools needed to analyze complex experimental data. Building such too

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

UAVFF3D: A Geometry-Aware Benchmark for Feed-Forward UAV 3D Reconstruction

DGX agent

arXiv:2605.17942v1 Announce Type: new Abstract: Feed-forward 3D reconstruction has recently demonstrated strong generalization across diverse scenes, yet its performance in UAV imagery remains underex

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

DGX agent

arXiv:2601.12696v2 Announce Type: replace Abstract: Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerab

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation

DGX agent

arXiv:2605.17140v1 Announce Type: cross Abstract: Brain tumor diagnosis is largely dependent on Magnetic Resonance Imaging (MRI) evaluation, which requires radiologists to synthesize thousands of imag

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

UniER: A Unified Benchmark for Item-level and Path-level Exercise Recommendation

DGX agent

arXiv:2605.16750v1 Announce Type: cross Abstract: Personalized exercise recommendation dynamically aligns pedagogical resources with individual knowledge mastery, which is crucial for satisfying stude

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings

DGX agent

arXiv:2605.17356v1 Announce Type: new Abstract: Existing works typically focus on presentation generation under isolated input settings, whereas real-world use cases span diverse scenarios, including

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

Universal Dynamics of Punctuated Progress

DGX agent

arXiv:2605.16719v1 Announce Type: cross Abstract: Scientific and technological frontiers advance through punctuated dynamics, yet the principles governing these dynamics remain poorly understood. Here

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Usenix'23 Extended Version: Smart Learning to Find Dumb Contracts

DGX agent

arXiv:2304.10726v3 Announce Type: replace-cross Abstract: We introduce the Deep Learning Vulnerability Analyzer (DLVA) for Ethereum smart contracts based on neural networks. We train DLVA to judge byt

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

UTOPYA: A Multimodal Deep Learning Framework for Physics-Informed Anomaly Detection and Time-Series Prediction

DGX agent

arXiv:2605.18188v1 Announce Type: new Abstract: Anomaly detection in batch processes is hindered by transient dynamics, scarce fault labels, and reliance on single-modality sensor data. This work intr

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

UVTran: Accurate Hole-Filling Parameterization with Transformers

DGX agent

arXiv:2605.16306v1 Announce Type: cross Abstract: In industrial design, N-sided hole filling is typically formulated as the construction of a single trimmed B-spline surface by minimizing a fairness e

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification

DGX agent

arXiv:2605.17691v1 Announce Type: cross Abstract: Automating the classification of negative treatment in legal precedent is a critical yet nuanced NLP task where misclassification carries significant

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Verifier-Guided Code Translation via Meta-Step Decoding

DGX agent

arXiv:2605.17626v1 Announce Type: new Abstract: Test-time scaling is an important mechanism for improving large language models, especially on tasks with deterministic verifiers. Code translation is a

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

Verify-Gated Completion as Admission Control in a Governed Multi-Agent Runtime: A Bounded Architecture Case Study

DGX agent

arXiv:2605.17998v1 Announce Type: cross Abstract: As multi-agent systems move from short interactions to tool-using workflows with specialized roles and persistent state, completion becomes a runtime-

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

DGX agent

arXiv:2605.17467v1 Announce Type: new Abstract: Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliabil

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

VGGT-CD: Training-Free Robust Registration for 3D Change Detection

DGX agent

arXiv:2605.16859v1 Announce Type: cross Abstract: 3D change detection from multi-view images is essential for urban monitoring, disaster assessment, and autonomous driving. However, existing methods p

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction

DGX agent

arXiv:2605.16911v1 Announce Type: new Abstract: 3D semantic occupancy prediction requires accurate 2D-to-3D feature lifting, yet current methods restrict camera geometry to initial projections. Subseq

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

[video] why we need a new continuity layer for long-running agents (claude did this video! all except the voice which was @elevenlabs)

DGX agent

This video discusses the architectural need for a continuity layer in long-running AI agents, explaining how agents require persistent memory and state management mechanisms to maintain coherence acro

model-releasesyohei-nakajima--x
19 May 2026
Model Releases

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models

DGX agent

arXiv:2605.18160v1 Announce Type: cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrati

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

DGX agent

arXiv:2602.04802v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing bench

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval

DGX agent

arXiv:2605.16481v1 Announce Type: cross Abstract: Long video understanding requires more than large context windows. It also needs a memory mechanism that decides what visual evidence to retain, keeps

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs

DGX agent

arXiv:2605.18172v1 Announce Type: new Abstract: Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Wasserstein Equilibrium Decoding for Reliable Medical Visual Question Answering

DGX agent

arXiv:2605.18313v1 Announce Type: cross Abstract: Small vision-language models (2-8B) are well-suited for clin- ical deployment due to privacy constraints, limited connectivity, and low-latency requir

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

DGX agent

arXiv:2601.06943v2 Announce Type: replace-cross Abstract: In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed ac

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

WavFlow: Audio Generation in Waveform Space

DGX agent

arXiv:2605.18749v1 Announce Type: cross Abstract: Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this wo

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong

DGX agent

arXiv:2509.22510v3 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) is the ability to satisfy desired objectives during generation, which is critical for trustworthy deployme

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

We were able to sit down with the @GoogleDeepmind team behind the new Gemini Omni Flash model to hear all of their behind-the-scenes stories…

DGX agent

We were able to sit down with the @GoogleDeepmind team behind the new Gemini Omni Flash model to hear all of their behind-the-scenes stories, memorable moments, and many, many (occasionally embarrassi

model-releasesgoogle-ai--x
19 May 2026
Model Releases

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games

DGX agent

arXiv:2605.17637v1 Announce Type: new Abstract: Coding agents are increasingly used as application builders, yet many evaluations still focus on source code, repository-level tests, or intermediate tr

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale

DGX agent

arXiv:2510.16252v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for web agents demands environments that are both effective for evaluation and efficient enough for large-scale on

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Weighted Flow Matching and Physics-Informed Nonlinear Filtering for Parameter Estimation in Digital Twins

DGX agent

arXiv:2605.17146v1 Announce Type: cross Abstract: Digital twins (DTs) rely on continuous synchronization between physical systems and their virtual counterparts through online parameter estimation und

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

WELD: The First Naturalistic Long-Period Small-Team Workplace Emotion Dataset for Ubiquitous Affective Computing

DGX agent

arXiv:2510.15221v2 Announce Type: replace Abstract: Affective computing has matured rapidly in laboratory settings, yet no prior dataset combines (i) months-to-years of duration, (ii) a naturalistic w

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

We’re adding new ways for people to identify AI-generated images and understand where they came from. In addition to C2PA Content Credential…

DGX agent

We’re adding new ways for people to identify AI-generated images and understand where they came from. In addition to C2PA Content Credentials, images now also contain a SynthID watermark, and can be i

model-releasesopenai--x
19 May 2026
Model Releases

We’re releasing Nemotron-Labs-Diffusion - the first Tri-mode LM family (3B/8B/14B) that switches between 1⃣Autoregressive, 2⃣Diffusion, and …

DGX agent

We’re releasing Nemotron-Labs-Diffusion - the first Tri-mode LM family (3B/8B/14B) that switches between 1⃣Autoregressive, 2⃣Diffusion, and 3⃣Self-Speculation decoding by simply changing the attention

model-releasesemad-mostaque--x
19 May 2026
Model Releases

What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models

DGX agent

arXiv:2605.18738v1 Announce Type: new Abstract: Medicine is inherently pluralistic. Principles such as autonomy, beneficence, nonmaleficence, and justice routinely conflict, and such ethical dilemmas

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

What Google I/O '26 means for developing agents on Google Cloud

DGX agent

At Google I/O, we introduced a unified development toolkit featuring Antigravity 2.0 and the Managed Agents API, giving developers better ways to build locally and deploy securely to the cloud on a sh

model-releasesgoogle-cloud-ai
19 May 2026
Model Releases

When Accuracy Is Not Enough: Uncertainty Collapse between Noisy Label Learning and Out-of-Distribution Detection

DGX agent

arXiv:2605.17795v1 Announce Type: cross Abstract: Learning with noisy labels (LNL) is typically benchmarked by closed-set classification accuracy, yet deployment often requires classifiers to reject o

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

When AI Tells You What You Want to Hear: Sycophantic Behavior of Large Language Models in Dementia Care Settings

DGX agent

arXiv:2605.16288v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in clinical and care settings. This exploratory study investigates whether LLMs exhibit sycophantic

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State

DGX agent

arXiv:2605.18580v1 Announce Type: new Abstract: Outcome-only evaluation can certify economically unsafe agents: a policy can hit a business KPI while violating deployable behavioral discipline. In hot

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

DGX agent

arXiv:2601.17887v2 Announce Type: replace Abstract: Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized ag

model-releasesarxiv-cs-ai
19 May 2026
← Previous
1…301302303304305…472
Next →