AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

Content type
87,042Total entries
1Added by human
87,041Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,106 results
Model Releases

High-Layer Attention Pruning with Rescaling

DGX agent

arXiv:2507.01900v3 Announce Type: replace Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation

DGX agent

arXiv:2608.07498v1 Announce Type: cross Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deploym

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

DGX agent

arXiv:2503.19990v4 Announce Type: replace Abstract: Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial reasoning

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

DGX agent

arXiv:2607.15509v2 Announce Type: replace-cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture desig

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks

DGX agent

arXiv:2507.03162v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These mod

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

DGX agent

arXiv:2608.09046v1 Announce Type: new Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure doe

model-releasesarxiv-cs-cl
11 Aug 2026
Research

MixFormer: Linear Transformer with Mixture of Memory Experts

DGX agent

arXiv:2608.09468v1 Announce Type: cross Abstract: State Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transformers in l

researcharxiv-cs-ai
11 Aug 2026
Model Releases

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

DGX agent

arXiv:2608.07993v1 Announce Type: new Abstract: Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor moti

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2510.18940v2 Announce Type: replace-cross Abstract: Existing parameter-efficient fine-tuning (PEFT) methods primarily fall into two categories: addition-based and selective in-situ adaptation. T

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Neurosymbolic Discovery of Algebraic Graph Constructions

DGX agent

arXiv:2608.08118v1 Announce Type: new Abstract: There are several methods for searching for graphs with prescribed properties, such as SAT solvers and specialized generators. These methods return the

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents

DGX agent

arXiv:2608.08264v1 Announce Type: new Abstract: Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift creates a new

model-releasesarxiv-cs-ai
11 Aug 2026
Research

On the Robustness of LLMs' Internal Representation of Code Correctness

DGX agent

arXiv:2608.08266v1 Announce Type: cross Abstract: Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no surprise, as r

researcharxiv-cs-ai
11 Aug 2026
Research

Optimal Learning Under Tsybakov Noise

DGX agent

arXiv:2608.08416v1 Announce Type: new Abstract: Probably Approximately Correct (PAC) learning [Val84] is a fundamental learning model that has been extensively investigated. In this model, H subseteq

researcharxiv-cs-lg
11 Aug 2026
Model Releases

P^{3}: Joint Program-and-Proof Planning for Verified Code Generation

DGX agent

arXiv:2608.09277v1 Announce Type: new Abstract: Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the program meets a

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

DGX agent

arXiv:2608.07631v1 Announce Type: cross Abstract: LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Preserving Item Semantics for Free: Rethinking Token Initialization in LLM-Based Generative Recommendation

DGX agent

arXiv:2608.07816v1 Announce Type: cross Abstract: Recent advances in generative recommendation (GR) leverage large language models (LLMs) as recommender backbones, enabling LLMs to directly generate r

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

DGX agent

arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on e

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation

DGX agent

arXiv:2503.22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction

DGX agent

arXiv:2608.09182v1 Announce Type: cross Abstract: Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Existing localiz

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning

DGX agent

arXiv:2608.09123v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a s

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough

DGX agent

arXiv:2608.07583v1 Announce Type: cross Abstract: Multi-agent LLM systems route among model-backed advisors, yet a deployer rarely knows before shipping whether routing will help at all. Prevailing ro

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

DGX agent

arXiv:2608.09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance,

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision

DGX agent

arXiv:2608.09097v1 Announce Type: new Abstract: Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

DGX agent

arXiv:2608.09771v1 Announce Type: new Abstract: Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

DGX agent

arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding

DGX agent

arXiv:2608.07915v1 Announce Type: new Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Th

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

DGX agent

arXiv:2608.09538v1 Announce Type: cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law

DGX agent

arXiv:2608.09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the ap

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

DGX agent

arXiv:2608.07617v1 Announce Type: new Abstract: Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatche

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction

DGX agent

arXiv:2608.07517v1 Announce Type: cross Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aw

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games

DGX agent

arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move fro

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

DGX agent

arXiv:2608.08795v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external o

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following

DGX agent

arXiv:2608.09154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly expected to follow long lists of constraints in complex instructions, and synthesizing instructions from a

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging

DGX agent

arXiv:2511.18121v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visu

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment

DGX agent

arXiv:2608.09892v1 Announce Type: new Abstract: Robot policy evaluation and deployment remain fragmented by model-specific software dependencies, data representations, and runtime interfaces, so that

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

Your Prompt Is Not the Only Prompt: How Much Do LLMs Weight Structured-Output Schema Descriptions?

DGX agent

arXiv:2608.08254v1 Announce Type: new Abstract: Structured output, where an LLM populates a predefined JSON schema, has become a default mechanism for data labeling and information extraction, but it

model-releasesarxiv-cs-ai
11 Aug 2026
Local Ai

ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling

DGX agent

arXiv:2608.07974v1 Announce Type: new Abstract: Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while preserving privacy. Although existing studies propos

local-aiarxiv-cs-lg
11 Aug 2026
Model Releases

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

DGX agent

arXiv:2608.07169v1 Announce Type: new Abstract: Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which strug

model-releasesarxiv-cs-ai
10 Aug 2026
Agents

An End-to-End Agent Auditing Engine

DGX agent

arXiv:2608.07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of d

agentsarxiv-cs-ai
10 Aug 2026
Model Releases

AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation

DGX agent

arXiv:2603.20986v2 Announce Type: replace Abstract: Phase-field modeling links thermodynamics and kinetics to microstructural evolution, but multiphysics frameworks such as MOOSE require expertise to

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

DGX agent

arXiv:2608.06930v1 Announce Type: new Abstract: Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation

DGX agent

arXiv:2608.06751v1 Announce Type: cross Abstract: Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution

DGX agent

arXiv:2608.06811v1 Announce Type: cross Abstract: Resolving a real software issue with a large language model (LLM) agent is a long repair episode, often tens to hundreds of steps spanning exploration

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training

DGX agent

arXiv:2608.06471v1 Announce Type: cross Abstract: Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in real-world s

model-releasesarxiv-cs-ai
10 Aug 2026
Local Ai

FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding

DGX agent

arXiv:2608.06819v1 Announce Type: cross Abstract: Token-level collaboration allows a large language model (LLM) to assist a small language model (SLM) when their predictions diverge. Existing methods

local-aiarxiv-cs-ai
10 Aug 2026
Local Ai

Human-AI Perceptual Alignment by Playing Hues and Cues

DGX agent

arXiv:2608.07141v1 Announce Type: new Abstract: Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks tha

local-aiarxiv-cs-cv
10 Aug 2026
Model Releases

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

DGX agent

arXiv:2608.07417v1 Announce Type: cross Abstract: Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents

DGX agent

arXiv:2608.06735v1 Announce Type: new Abstract: Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as

model-releasesarxiv-cs-ai
10 Aug 2026
← Previous
1…338339340341342…1065
Next →