AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
91,060Total entries
1Added by human
91,059Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,793 results
Research

On the Robustness of LLMs' Internal Representation of Code Correctness

DGX agent

arXiv:2608.08266v1 Announce Type: cross Abstract: Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no surprise, as r

researcharxiv-cs-ai
11 Aug 2026
Research

Optimal Learning Under Tsybakov Noise

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2608.08416v1 Announce Type: new Abstract: Probably Approximately Correct (PAC) learning [Val84] is a fundamental learning model that has been extensively investigated. In this model, H subseteq

researcharxiv-cs-lg
11 Aug 2026
Model Releases

P^{3}: Joint Program-and-Proof Planning for Verified Code Generation

DGX agent

arXiv:2608.09277v1 Announce Type: new Abstract: Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the program meets a

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

DGX agent

arXiv:2608.07631v1 Announce Type: cross Abstract: LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Preserving Item Semantics for Free: Rethinking Token Initialization in LLM-Based Generative Recommendation

DGX agent

arXiv:2608.07816v1 Announce Type: cross Abstract: Recent advances in generative recommendation (GR) leverage large language models (LLMs) as recommender backbones, enabling LLMs to directly generate r

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

DGX agent

arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on e

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation

DGX agent

arXiv:2503.22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction

DGX agent

arXiv:2608.09182v1 Announce Type: cross Abstract: Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Existing localiz

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning

DGX agent

arXiv:2608.09123v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a s

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough

DGX agent

arXiv:2608.07583v1 Announce Type: cross Abstract: Multi-agent LLM systems route among model-backed advisors, yet a deployer rarely knows before shipping whether routing will help at all. Prevailing ro

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

DGX agent

arXiv:2608.09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance,

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision

DGX agent

arXiv:2608.09097v1 Announce Type: new Abstract: Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

DGX agent

arXiv:2608.09771v1 Announce Type: new Abstract: Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

DGX agent

arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding

DGX agent

arXiv:2608.07915v1 Announce Type: new Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Th

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

DGX agent

arXiv:2608.09538v1 Announce Type: cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law

DGX agent

arXiv:2608.09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the ap

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

DGX agent

arXiv:2608.07617v1 Announce Type: new Abstract: Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatche

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction

DGX agent

arXiv:2608.07517v1 Announce Type: cross Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aw

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games

DGX agent

arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move fro

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

DGX agent

arXiv:2608.08795v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external o

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following

DGX agent

arXiv:2608.09154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly expected to follow long lists of constraints in complex instructions, and synthesizing instructions from a

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging

DGX agent

arXiv:2511.18121v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visu

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment

DGX agent

arXiv:2608.09892v1 Announce Type: new Abstract: Robot policy evaluation and deployment remain fragmented by model-specific software dependencies, data representations, and runtime interfaces, so that

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

Your Prompt Is Not the Only Prompt: How Much Do LLMs Weight Structured-Output Schema Descriptions?

DGX agent

arXiv:2608.08254v1 Announce Type: new Abstract: Structured output, where an LLM populates a predefined JSON schema, has become a default mechanism for data labeling and information extraction, but it

model-releasesarxiv-cs-ai
11 Aug 2026
Local Ai

ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling

DGX agent

arXiv:2608.07974v1 Announce Type: new Abstract: Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while preserving privacy. Although existing studies propos

local-aiarxiv-cs-lg
11 Aug 2026
Model Releases

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

DGX agent

arXiv:2608.07169v1 Announce Type: new Abstract: Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which strug

model-releasesarxiv-cs-ai
10 Aug 2026
Agents

An End-to-End Agent Auditing Engine

DGX agent

arXiv:2608.07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of d

agentsarxiv-cs-ai
10 Aug 2026
Model Releases

AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation

DGX agent

arXiv:2603.20986v2 Announce Type: replace Abstract: Phase-field modeling links thermodynamics and kinetics to microstructural evolution, but multiphysics frameworks such as MOOSE require expertise to

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

DGX agent

arXiv:2608.06930v1 Announce Type: new Abstract: Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation

DGX agent

arXiv:2608.06751v1 Announce Type: cross Abstract: Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution

DGX agent

arXiv:2608.06811v1 Announce Type: cross Abstract: Resolving a real software issue with a large language model (LLM) agent is a long repair episode, often tens to hundreds of steps spanning exploration

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training

DGX agent

arXiv:2608.06471v1 Announce Type: cross Abstract: Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in real-world s

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks

DGX agent

Having a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti

model-releasesr-localllama
10 Aug 2026
Local Ai

FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding

DGX agent

arXiv:2608.06819v1 Announce Type: cross Abstract: Token-level collaboration allows a large language model (LLM) to assist a small language model (SLM) when their predictions diverge. Existing methods

local-aiarxiv-cs-ai
10 Aug 2026
Model Releases

Google named a Leader in The Forrester Wave™: AI Platforms, Q3 2026

DGX agent

At Google Cloud, we help organizations of all sizes build and operationalize complex agentic workflows with total confidence. By combining world-class AI research with an open, fully integrated AI pla

model-releasesgoogle-cloud-ai
10 Aug 2026
Local Ai

Human-AI Perceptual Alignment by Playing Hues and Cues

DGX agent

arXiv:2608.07141v1 Announce Type: new Abstract: Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks tha

local-aiarxiv-cs-cv
10 Aug 2026
Model Releases

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

DGX agent

arXiv:2608.07417v1 Announce Type: cross Abstract: Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents

DGX agent

arXiv:2608.06735v1 Announce Type: new Abstract: Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face

DGX agent

Looks like the Ling team open weighted a much smaller version of the Ling-3.0-flash they open weighted a few days ago. It's 8B params with 1.3B active, and seems to fall between the 4B and 8-12B Qwen

model-releasesr-localllama
10 Aug 2026
Tutorials

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

DGX agent

arXiv:2608.06411v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of

tutorialsarxiv-cs-ai
10 Aug 2026
Model Releases

Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework

DGX agent

arXiv:2608.06909v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, a

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Modular TTT: Rethinking Test-Time Training as Composable Modules

DGX agent

arXiv:2608.07110v1 Announce Type: cross Abstract: Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite

model-releasesarxiv-cs-cl
10 Aug 2026
Model Releases

MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs

DGX agent

arXiv:2607.08970v2 Announce Type: replace-cross Abstract: Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observa

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks

DGX agent

arXiv:2505.04397v3 Announce Type: replace-cross Abstract: Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexplore

model-releasesarxiv-cs-ai
10 Aug 2026
Safety

Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation

DGX agent

arXiv:2608.07176v1 Announce Type: cross Abstract: Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training

safetyarxiv-cs-ai
10 Aug 2026
Safety

Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation

DGX agent

arXiv:2608.07154v1 Announce Type: cross Abstract: Open-source robotics and foundation models have lowered the barrier to embodied AI, yet language-guided laboratory automation still requires reliable

safetyarxiv-cs-ai
10 Aug 2026
Model Releases

RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs

DGX agent

arXiv:2608.07088v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing trai

model-releasesarxiv-cs-ai
10 Aug 2026
← Previous
1…436437438439440…1371
Next →