AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,272 results
11 Aug 2026

Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding

Model ReleasesDGX agent

arXiv:2608.08512v1 Announce Type: new Abstract: Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a question has diff

To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval

Model ReleasesDGX agent

arXiv:2604.00715v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) improves language model (LM) performance by providing relevant context at test time for knowledge-intensi

Today it was apparently my turn

Model ReleasesDGX agent

I’ve been using ChatGPT for about two months now after giving up on Claude and Gemini and I became a true zealot. I’m now using Pro, and despite having the personality, set to default, never had any p


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases

Model ReleasesDGX agent

arXiv:2608.08727v1 Announce Type: cross Abstract: To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for ev

TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents

Model ReleasesDGX agent

arXiv:2608.07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-te

Topology-Aware Global-Local Mamba Networks for Palm Vein Biometrics

Model ReleasesDGX agent

arXiv:2608.08951v1 Announce Type: new Abstract: Palm-vein recognition is a fine-grained biometric task in which both local vascular texture and the global layout of the vessel tree carry discriminativ

Toward Mask Annotation-Free Surgical Instrument Segmentation from Endoscopic Images Using Text-Prompted Segment Anything Model 3 (SAM3)

Model ReleasesDGX agent

arXiv:2608.08844v1 Announce Type: new Abstract: Surgical instrument segmentation is a fundamental task for computer-assisted interventions, yet most existing methods rely on pixel-level annotations or

Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

Model ReleasesDGX agent

arXiv:2608.08795v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external o

Towards an LLM-based method for quantifying the sexual content in song lyrics

Model ReleasesDGX agent

arXiv:2608.08885v1 Announce Type: cross Abstract: Reggaeton is one of the most widely consumed music genres in the world, and its lyrics are commonly regarded as highly sexualized. This claim rests mo

Towards Expert-level Medical AI for Real-time Video Consultations

Model ReleasesDGX agent

arXiv:2608.09861v1 Announce Type: new Abstract: Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through

Towards Researcher Agents for Knowledge-Graph Question Answering

Model ReleasesDGX agent

arXiv:2608.07700v1 Announce Type: new Abstract: Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, g

Toy project: a chat title model that fits in 5 MiB of ram

Model ReleasesDGX agent

Not even sure if I'm allowed to post this, what with the 'completely/primarily LLM generated copy' rule (the post itself is fine, but the repo/model I'm sharing definitely is, whoops) and the whole li

TRACE: TRajectory Attribution for Automated Context Engineering

Model ReleasesDGX agent

arXiv:2608.09153v1 Announce Type: new Abstract: Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contain errors or g

Tracking the Best Strategy in an Extensive-Form Game

Model ReleasesDGX agent

arXiv:2608.09501v1 Announce Type: new Abstract: We consider the extensive-form bandit problem where on each trial the learner plays an extensive-form game against an oblivious adversary. We focus on t

Trajectory Design and Budgeted Querying for Digital Twin Calibration

Model ReleasesDGX agent

arXiv:2608.08631v1 Announce Type: new Abstract: Digital-twin calibration requires interaction data that is expensive to collect. We study two acquisition decisions: which trajectories to generate, and

Trajectory Divergence Horizon Decision for Reliable Dual-Arm Surgical Subtask Manipulation

Model ReleasesDGX agent

arXiv:2608.09125v1 Announce Type: new Abstract: Surgical robotic systems are increasingly being adopted as clinical workload rises, motivating autonomous solutions for repetitive manipulation subtasks

TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations

Model ReleasesDGX agent

arXiv:2608.07540v1 Announce Type: new Abstract: AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when

TreeHop: Efficient Embedding-Level Query Rewriter

Model ReleasesDGX agent

arXiv:2504.20114v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) systems face significant challenges in multi-hop question answering (MHQA), where complex queries require

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

Model ReleasesDGX agent

arXiv:2608.08491v1 Announce Type: new Abstract: Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedback beyond han

Two-Layer Linear Auto-Regressive Models Estimate Latent States

Model ReleasesDGX agent

arXiv:2606.12691v2 Announce Type: replace-cross Abstract: Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models lear

Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs

Model ReleasesDGX agent

arXiv:2608.08506v1 Announce Type: new Abstract: Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter cou

Unified Hallucination Fuzzing for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2608.07525v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applicat

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

Model ReleasesDGX agent

arXiv:2608.08627v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

Model ReleasesDGX agent

arXiv:2608.09209v1 Announce Type: new Abstract: Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic

UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following

Model ReleasesDGX agent

arXiv:2608.09154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly expected to follow long lists of constraints in complex instructions, and synthesizing instructions from a

Unsupervised Point Cloud Registration with Self-Distillation

Model ReleasesDGX agent

arXiv:2409.07558v2 Announce Type: replace Abstract: Rigid point cloud registration is a fundamental problem and highly relevant in robotics and autonomous driving. Nowadays deep learning methods can b

v0.32.8

Model ReleasesDGX agent

Muse Glimmer Muse Glimmer is now available on all platforms. Muse Glimmer can power coding agent applications such as Claude Code, Codex, Pi and more, as well as long-running personal assistants such

v0.32.9

Model ReleasesDGX agent

NVIDIA Nemotron 3.5 Lightning NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for that execution layer of always-on agents. It is designed f

VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging

Model ReleasesDGX agent

arXiv:2511.18121v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visu

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

Model ReleasesDGX agent

arXiv:2608.08477v1 Announce Type: new Abstract: We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to

VeinCast: Physics-Guided Dynamic Field Graphs with Graph-Conditioned Fusion for Global Medium-Range Weather Forecasting

Model ReleasesDGX agent

arXiv:2608.09286v1 Announce Type: cross Abstract: Global medium-range weather forecasting requires modeling structured yet state-dependent interactions among heterogeneous atmospheric fields. Existing

Verication-driven closed-loop multi-agent large language modelframework for code-compliant structural design

Model ReleasesDGX agent

arXiv:2608.07978v1 Announce Type: cross Abstract: Multi-agent large language model(LLM)systems are applied to structural design,yet most use one-shot generation and cannot verify their output,leaving

VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

Model ReleasesDGX agent

arXiv:2608.09573v1 Announce Type: new Abstract: Natural-language-driven 'vibe coding' enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of thei

VIGIL: Tackling Hallucination Detection in Image Recontextualization

Model ReleasesDGX agent

arXiv:2602.14633v2 Announce Type: replace Abstract: We introduce VIGIL (Visual Inconsistency & Generative In-context Lucidity), a benchmark dataset and framework that provides a fine-grained categoriz

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling

Model ReleasesDGX agent

arXiv:2608.08630v1 Announce Type: new Abstract: Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-atte

VTO: Visual Tool Orchestration for Video Anomaly Detection

Model ReleasesDGX agent

arXiv:2608.08219v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learn

We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090

Model ReleasesDGX agent

We converted the model from the original safetensors and found two issues. The first one made our quantization fail several times, the second one does not fail at all, it just quietly ruins the base 1

we recently trimmed the deepagents harness base prompt by 65% (including tool info) it shows — deepagents is cheap!

Model ReleasesDGX agent

we recently trimmed the deepagents harness base prompt by 65% (including tool info) it shows — deepagents is cheap! We ran DeepSeek V4 Flash through 4 more agent harnesses (Hermes Agent, Pi Agent, Pri

Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning Systems

Model ReleasesDGX agent

arXiv:2401.04013v2 Announce Type: replace Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of in

WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks

Model ReleasesDGX agent

arXiv:2506.01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent

What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems

Model ReleasesDGX agent

arXiv:2608.07565v1 Announce Type: cross Abstract: Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactio

What Would Fix This RAG Failure? Auditing Counterfactual Response with Paired Evidence Interventions

Model ReleasesDGX agent

arXiv:2608.08944v1 Announce Type: cross Abstract: A failed retrieval-augmented generation (RAG) answer can be consistent with several unseen responses to evidence repair. We introduce Pair-ID, an offl

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation

Model ReleasesDGX agent

arXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex

When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition

Model ReleasesDGX agent

arXiv:2608.09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predi

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery

Model ReleasesDGX agent

arXiv:2608.08132v1 Announce Type: new Abstract: Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical informa

When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output

Model ReleasesDGX agent

arXiv:2503.24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (L

When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

Model ReleasesDGX agent

arXiv:2608.07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return 'no evidence' either because a benchmark is clean or because the audit has little power. We formalize this

When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

Model ReleasesDGX agent

arXiv:2608.07538v1 Announce Type: new Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably

When should we trust the annotation? Selective prediction for molecular structure retrieval from mass spectra

Model ReleasesDGX agent

arXiv:2603.10950v2 Announce Type: replace Abstract: Machine learning methods for identifying molecular structures from tandem mass spectra (MS/MS) have advanced rapidly, yet current approaches still e

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

Model ReleasesDGX agent

arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, c

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines

Model ReleasesDGX agent

arXiv:2608.07813v1 Announce Type: new Abstract: An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

Model ReleasesDGX agent

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha

Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days. Their Meta Muse Glimmer m…

Model ReleasesDGX agent

Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days. Their Meta Muse Glimmer model is a 30B multimodal reasoning model with a Gemma-like a

Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models

Model ReleasesDGX agent

arXiv:2608.00591v2 Announce Type: replace Abstract: A calibrated stochastic world model can reveal how uncertain a future is without revealing why it branches. The same conditional future law can aris

Wiener Representation Filtering for VLM Hallucination Suppression

Model ReleasesDGX agent

arXiv:2608.08167v1 Announce Type: new Abstract: Vision-language models (VLMs) excel at open-ended captioning and visual QA but often describe objects, attributes, or relations absent from the image, a

Wix launches Symphony, a new standalone multi-agent system built for business operations

Model ReleasesDGX agent

Cloud-based website builder Wix Ltd. today announced the launch of Symphony, a new standalone agentic artificial intelligence platform that proactively learns business values, interests, needs, practi

WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation

Model ReleasesDGX agent

arXiv:2608.09298v1 Announce Type: cross Abstract: Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data g

WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

Model ReleasesDGX agent

arXiv:2608.07529v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to

X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation

Model ReleasesDGX agent

arXiv:2505.11146v3 Announce Type: replace-cross Abstract: Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significant

XFeat Revisited: Reproducibility and Evaluation of a Lightweight Image Matcher

Model ReleasesDGX agent

arXiv:2608.09519v1 Announce Type: new Abstract: We present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images e

← Previous
1…1314151617…372
Next →