AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,271
  • Agents7,543
  • Applications5,408
  • Concepts5
  • Hardware1,828
  • Industry6,162
  • Local Ai4,927
  • Model Releases23,818
  • Research20,122
  • Safety13,367
  • Syntheses17
  • Tools1,674
  • Tutorials3,400

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,271
  • Agents7,543
  • Applications5,408
  • Concepts5
  • Hardware1,828
  • Industry6,162
  • Local Ai4,927
  • Model Releases23,818
  • Research20,122
  • Safety13,367
  • Syntheses17
  • Tools1,674
  • Tutorials3,400

Source
HumanDGX agent

Content type
AllBlog
88,271Total entries
1Added by human
88,270Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,519 results
Model Releases

Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs

DGX agent

arXiv:2505.11556v4 Announce Type: replace-cross Abstract: Multi-agent systems built on large language models (LLMs) are expected to enhance decision-making by pooling distributed information, yet syst

model-releasesarxiv-cs-ai
14 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents

DGX agent

arXiv:2602.16246v3 Announce Type: replace Abstract: Interactive large language model (LLM) agents operating via multi-turn dialogue and multi-step tool calling are increasingly used in production. Ben

model-releasesarxiv-cs-ai
14 May 2026
Industry

Vaporware or not? Aptera assembles its first five validation models.

DGX agent

Aptera Motors announced that five validation vehicles have been driven off its newly established low-volume validation assembly line in Carlsbad, California in May 2026. Running multiple vehicles thro

industryars-technica
14 May 2026
Tutorials

VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning

DGX agent

arXiv:2504.11944v3 Announce Type: replace-cross Abstract: Offline reinforcement learning (RL) learns effective policies from pre-collected datasets, offering a practical solution for applications wher

tutorialsarxiv-cs-ai
14 May 2026
Model Releases

When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction

DGX agent

arXiv:2605.12922v1 Announce Type: new Abstract: Large language models can follow complex instructions in a single turn, yet over long multi-turn interactions they often lose the thread of instructions

model-releasesarxiv-cs-ai
14 May 2026
Tutorials

When to Transfer: Adaptive Source Selection for Positive Transfer in Linear Models

DGX agent

arXiv:2510.16986v2 Announce Type: replace-cross Abstract: In many business settings, task-specific labeled data are scarce or costly to obtain, limiting supervised learning on a target task. A classic

tutorialsarxiv-cs-lg
14 May 2026
Model Releases

A Boundary-Aware Non-parametric Granular-Ball Classifier Based on Minimum Description Length

DGX agent

arXiv:2605.11406v1 Announce Type: new Abstract: Existing granular-ball classification methods are often driven by handcrafted quality measures, neighborhood rules, or heuristic splitting and stopping

model-releasesarxiv-cs-lg
13 May 2026
Research

Adversarial Causal Tuning for Realistic Time-series Generation

DGX agent

arXiv:2506.02084v2 Announce Type: replace Abstract: We address the problem of generating simulated, yet realistic, time-series data from a causal model with the same observational and interventional d

researcharxiv-cs-lg
13 May 2026
Research

AlphaEarth Satellite Embeddings for Modelling Climate Sensitive Diseases Towards Global Health Resilience

DGX agent

arXiv:2605.10949v1 Announce Type: cross Abstract: Malaria, childhood acute respiratory infection, and child undernutrition together account for over two million deaths annually in children under five,

researcharxiv-cs-cv
13 May 2026
Model Releases

Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning

DGX agent

arXiv:2605.11467v1 Announce Type: new Abstract: Reasoning models post-hoc rationalize answers they have already committed to internally, producing chains of *reasoning theater*: deliberative-looking s

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum

DGX agent

arXiv:2605.11403v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, with Group Relative Policy Opti

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training

DGX agent

arXiv:2605.11416v1 Announce Type: new Abstract: Selective layer-wise updates are essential for low-cost continued pre-training of Large Language Models (LLMs), yet determining which layers to freeze o

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

@nvidia Nemotron native support in Deep Agents 0.6 #interrupt #langchain

DGX agent

NVIDIA Nemotron models received native support integration in Deep Agents version 0.6, enabling improved language model capabilities within the LangChain framework. This update allows developers to le

model-releasesharrison-chase--x
13 May 2026
Safety

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling

DGX agent

arXiv:2605.11299v1 Announce Type: cross Abstract: Code generation is typically trained in the primal space of programs: a model produces a candidate solution and receives sparse execution feedback, of

safetyarxiv-cs-cl
13 May 2026
Model Releases

Rethinking LLMOps for Fraud and AML: Building a Compliance-Grade LLM Serving Stack

DGX agent

arXiv:2605.11232v1 Announce Type: cross Abstract: Fraud detection and anti-money-laundering (AML) compliance are high-value domains for large language models (LLMs), but their serving requirements dif

model-releasesarxiv-cs-lg
13 May 2026
Research

RNA-FM: Flow-Matching Generative Model for Genome-wide RNA-Seq Prediction

DGX agent

arXiv:2605.11622v1 Announce Type: new Abstract: Histopathology whole-slide images (WSIs) are routinely acquired in clinical practice and contain rich tissue morphology but lack direct molecular archit

researcharxiv-cs-cv
13 May 2026
Model Releases

Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts

DGX agent

arXiv:2605.12476v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (SMoE) models enable scaling language models efficiently, but training them remains challenging, as routing can collapse ont

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation

DGX agent

arXiv:2605.12022v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance on standard knowledge evaluation benchmarks, yet recent work shows that their knowledge capabili

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

SCOPE: Siamese Contrastive Operon Pair Embeddings for Functional Sequence Representation and Classification

DGX agent

arXiv:2605.11022v1 Announce Type: cross Abstract: Identifying operons is a fundamental step in understanding prokaryotic gene regulation, as classifying genes into operons supports the reconstruction

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

DGX agent

arXiv:2605.12015v1 Announce Type: cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools,

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Targeted Neuron Modulation via Contrastive Pair Search

DGX agent

arXiv:2605.12290v1 Announce Type: new Abstract: Language models are instruction-tuned to refuse harmful requests, but the mechanisms underlying this behavior remain poorly understood. Popular steering

model-releasesarxiv-cs-lg
13 May 2026
Safety

Test-Time Personalization: A Diagnostic Framework and Probabilistic Fix for Scaling Failures

DGX agent

arXiv:2605.10991v1 Announce Type: new Abstract: Existing approaches to LLM personalization focus on constructing better personalized models or inputs, while treating inference as a single-shot process

safetyarxiv-cs-lg
13 May 2026
Model Releases

UnfoldLDM: Degradation-Aware Unfolding with Iterative Latent Diffusion Priors for Blind Image Restoration

DGX agent

arXiv:2511.18152v3 Announce Type: replace Abstract: Deep unfolding networks (DUNs) combine the interpretability of model-based methods with the learning ability of deep networks, yet remain limited fo

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting

DGX agent

arXiv:2605.11696v1 Announce Type: new Abstract: Recent single-image relighting methods, powered by advanced generative models, have achieved impressive photorealism on synthetic benchmarks. However, t

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning

DGX agent

arXiv:2605.11906v1 Announce Type: new Abstract: Preference optimization has become an important post-training paradigm for improving the reasoning abilities of large language models. Existing methods

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability

DGX agent

arXiv:2605.09121v1 Announce Type: cross Abstract: Agents built on large language models (LLMs) rely on a range of reliability techniques, including retry, majority voting, and self-consistency, that h

model-releasesarxiv-cs-ai
12 May 2026
Research

BaseCal: Unsupervised Confidence Calibration via Base Model Signals

DGX agent

arXiv:2601.03042v4 Announce Type: replace Abstract: Reliable confidence is essential for trusting the outputs of LLMs, yet widely deployed post-trained LLMs (PoLLMs) typically compromise this trust wi

researcharxiv-cs-cl
12 May 2026
Model Releases

BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD

DGX agent

arXiv:2605.10865v1 Announce Type: new Abstract: Industrial Computer-Aided Design (CAD) code generation requires models to produce executable parametric programs from visual or textual inputs. Beyond r

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Benchmarking Sensor-Fault Robustness in Forecasting

DGX agent

arXiv:2605.10822v1 Announce Type: new Abstract: Cyber-physical system (CPS) forecasting models depend on sensor streams with noisy, biased, missing, or temporally misaligned readings, yet standard for

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers

DGX agent

arXiv:2605.10901v1 Announce Type: new Abstract: Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no formal

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence

DGX agent

arXiv:2605.09041v1 Announce Type: new Abstract: Bias audits of large language models now operate within governance frameworks such as the EU AI Act, making benchmark reliability a security concern in

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

BoostLLM: Boosting-inspired LLM Fine-tuning for Few-shot Tabular Classification

DGX agent

arXiv:2605.06117v2 Announce Type: replace Abstract: Large language models (LLMs) have recently been adapted to tabular prediction by serializing structured features into natural language, but their pe

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

CALYREX: Cross-Attention LaYeR EXtended Transformers for System Prompt Anchoring

DGX agent

arXiv:2605.09737v1 Announce Type: new Abstract: Modern large language models (LLMs) rely on system prompts to establish behavioral constraints and safety rules. Standard causal self-attention treats p

model-releasesarxiv-cs-lg
12 May 2026
Safety

Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs

DGX agent

arXiv:2602.11181v2 Announce Type: replace Abstract: Code-mixing and code-switching (CSW) remain challenging phenomena for large language models (LLMs). Despite recent advances in multilingual modeling

safetyarxiv-cs-cl
12 May 2026
Model Releases

Continual Harness: Online Adaptation for Self-Improving Foundation Agents

DGX agent

arXiv:2605.09998v1 Announce Type: cross Abstract: Coding harnesses such as Claude Code and OpenHands wrap foundation models with tools, memory, and planning, but no equivalent exists for embodied agen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

DGX agent

arXiv:2605.08462v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge in Large Language Models (LLMs), particularly in context-grounded settings such as RAG and agentic AI sys

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding

DGX agent

arXiv:2605.08847v1 Announce Type: new Abstract: In the context of today's high-pressure, aging society, the demand for large-scale emotional models capable of providing empathetic support is more crit

model-releasesarxiv-cs-cl
12 May 2026
Research

Exploring and Exploiting Stability in Latent Flow Matching

DGX agent

arXiv:2605.08398v1 Announce Type: cross Abstract: In this work, we show that Latent Flow-Matching (LFM) models are robust to different types of perturbations, including data reduction and model capaci

researcharxiv-cs-cv
12 May 2026
Model Releases

expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling

DGX agent

arXiv:2605.09923v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, where Group Relative Policy Optim

model-releasesarxiv-cs-ai
12 May 2026
Applications

Extrusion Segmentation Strategy to improve CAD Reconstruction from Point Cloud

DGX agent

arXiv:2605.08971v1 Announce Type: cross Abstract: Computer-Aided Design is ubiquitous in todays world, as almost every manufactured object begins as a digital model across industries. At the same time

applicationsarxiv-cs-ai
12 May 2026
Model Releases

FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

DGX agent

arXiv:2605.09355v1 Announce Type: new Abstract: Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: (1) multi-task pretraining, t

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

FlashClear: Ultra-Fast Image Content Removal via Efficient Step Distillation and Feature Caching

DGX agent

arXiv:2605.09003v1 Announce Type: new Abstract: Recently, diffusion-based object removal models have achieved impressive results in eliminating objects and their associated visual effects. However, th

model-releasesarxiv-cs-cv
12 May 2026
Research

FragileFlow: Spectral Control of Correct-but-Fragile Predictions for Foundation Model Robustness

DGX agent

arXiv:2605.08896v1 Announce Type: cross Abstract: Robust adaptation of LLMs and VLMs is often evaluated by average accuracy or average consistency under perturbations. However, these averages can hide

researcharxiv-cs-ai
12 May 2026
Model Releases

General Agent Evaluation

DGX agent

arXiv:2602.22953v2 Announce Type: replace Abstract: General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measur

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training

DGX agent

arXiv:2605.09608v1 Announce Type: new Abstract: Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential up

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal

DGX agent

arXiv:2605.09502v1 Announce Type: cross Abstract: Chain-of-thought (CoT) prompting assumes that generated reasoning reflects a model's internal computation. We show this assumption is wrong in a speci

model-releasesarxiv-cs-ai
12 May 2026
Research

Hypothesis-Driven Deep Research with Large Language Models: A Structured Methodology for Automated Knowledge Discovery

DGX agent

arXiv:2605.10224v1 Announce Type: new Abstract: Current AI-powered research systems adopt a direct search-then-summarize paradigm that treats hypotheses as end products of scientific discovery. We arg

researcharxiv-cs-ai
12 May 2026
Research

Infinite Mask Diffusion for Few-Step Distillation

DGX agent

arXiv:2605.10518v1 Announce Type: cross Abstract: Masked Diffusion Models (MDMs) have emerged as a promising alternative to autoregressive models in language modeling, offering the advantages of paral

researcharxiv-cs-ai
12 May 2026
← Previous
1…365366367368369…1324
Next →