AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,106 results
Model Releases

BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data

DGX agent

arXiv:2605.25549v1 Announce Type: cross Abstract: High-quality expert chain-of-thought (CoT) data is one of the core bottlenecks in large language model (LLM) post-training. Existing data production m

model-releasesarxiv-cs-ai
26 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

BODHI: Precise OS Kernel Specification Inference

DGX agent

arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express

DGX agent

arXiv:2605.25891v1 Announce Type: cross Abstract: We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference

DGX agent

arXiv:2510.02361v2 Announce Type: replace-cross Abstract: Transformer-based large models excel in natural language processing and computer vision, but face severe computational inefficiencies due to t

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

CITYREP: A Unified Benchmark for Urban Representations Across Cities, Tasks, and Modalities

DGX agent

arXiv:2605.26036v1 Announce Type: new Abstract: Urban representation learning encodes complex urban environments into general-purpose embeddings for diverse downstream tasks and emerging urban foundat

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

CMAP: Cross-Modal Adaptive Prompting for Multi-Domain Task-Incremental Learning

DGX agent

arXiv:2605.25708v1 Announce Type: cross Abstract: Multi-domain task-incremental learning requires a model to sequentially acquire knowledge across visually diverse domains without forgetting prior tas

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

DRInQ: Evaluating Conversational Implicature with Controlled Context Variation

DGX agent

arXiv:2605.24267v1 Announce Type: new Abstract: Human conversation relies heavily on conversational implicature, in which speakers convey meanings that are suggested rather than explicitly stated. Alt

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

DropoutTS: Sample-Adaptive Dropout for Robust Time Series Forecasting

DGX agent

arXiv:2601.21726v2 Announce Type: replace Abstract: Deep time series models are vulnerable to noisy data ubiquitous in real-world applications. Existing robustness strategies either prune data or rely

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Efficient DP-SGD for LLMs with Randomized Clipping

DGX agent

arXiv:2605.24879v1 Announce Type: new Abstract: Large language models (LLMs) are trained on vast datasets that may contain sensitive information. Differential privacy (DP), the de facto standard for f

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

DGX agent

arXiv:2605.25052v1 Announce Type: new Abstract: Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these t

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations

DGX agent

arXiv:2507.07644v4 Announce Type: replace Abstract: We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large language models (LLMs). FloorplanQA is grounded in struct

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap

DGX agent

arXiv:2605.24432v1 Announce Type: new Abstract: Large Language Model (LLM) interactions are typically underspecified, with users clarifying all necessary details across multiple conversational turns.

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

DGX agent

arXiv:2602.18527v2 Announce Type: replace-cross Abstract: Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. T

model-releasesarxiv-cs-ai
26 May 2026
Safety

Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation

DGX agent

arXiv:2605.25036v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) extend large language models with visual understanding, but remain vulnerable to hallucination, where outputs are

safetyarxiv-cs-ai
26 May 2026
Model Releases

Llamion Technical Report

DGX agent

arXiv:2605.25676v1 Announce Type: new Abstract: We release Llamion, a family of 14B-parameter open-weight language models obtained by transforming Orion-14B into the standardized Llama-family architec

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

DGX agent

arXiv:2505.23764v3 Announce Type: replace-cross Abstract: Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, h

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Mosaic: Compositional Multi-Concept Erasure via Vector Field Blending

DGX agent

arXiv:2605.25574v1 Announce Type: cross Abstract: Concept erasure has emerged as a key research direction for ensuring safe and ethical image synthesis in Text-to-Image (T2I) models. While existing st

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Omissive Bias in Religious Representation: Benchmarking LLM Answers to Everyday Ethical Decision-making

DGX agent

arXiv:2605.24319v1 Announce Type: new Abstract: As large language models become a default source of guidance on personal, moral, and existential questions, it matters whether they draw on the religiou

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Optimizing Digital Therapeutic Interventions: Online Learning under Endogenous Adherence

DGX agent

arXiv:2605.24261v1 Announce Type: new Abstract: A critical challenge facing clinicians managing chronic disease interventions is sustaining long-run patient health given limited information and resour

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Raon-Speech Technical Report

DGX agent

arXiv:2605.23912v1 Announce Type: cross Abstract: We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English and Korean speech understanding, answering, and gen

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection

DGX agent

arXiv:2605.24834v1 Announce Type: cross Abstract: Large language model (LLM) safety classifiers such as Llama Guard are effective at detecting overtly harmful prompts but remain vulnerable to adversar

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

RepetitionCurse: Measuring and Understanding Router Imbalance in Mixture-of-Experts LLMs under DoS Stress

DGX agent

arXiv:2512.23995v2 Announce Type: replace-cross Abstract: Mixture-of-Experts architectures have become the standard for scaling large language models due to their superior parameter efficiency. To acc

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Riemannian-Manifold Steering: Geometry-Aware Generative Autoencoders for Label-Free Steering

DGX agent

arXiv:2605.24942v1 Announce Type: cross Abstract: Steering a language model - intervening on its internal activations to change downstream behaviour - has recently expanded beyond linear interpolation

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

RL with Learnable Textual Feedback: A Bilevel Approach

DGX agent

arXiv:2605.24547v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards can improve LLM reasoning, but learning remains sample-inefficient when terminal rewards are sparse. This

model-releasesarxiv-cs-lg
26 May 2026
Safety

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

DGX agent

arXiv:2605.24817v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have become an increasingly important paradigm for scaling Large Language Models (LLMs). As MoE models are incr

safetyarxiv-cs-cl
26 May 2026
Model Releases

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

DGX agent

arXiv:2605.24117v1 Announce Type: new Abstract: Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience c

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs

DGX agent

arXiv:2605.25534v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at structural reasoning yet suffer from a sharp logical brittleness in structural consistency. We term th

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation

DGX agent

arXiv:2605.25488v1 Announce Type: cross Abstract: Audio-driven talking-head generation has achieved remarkable progress with recent models such as AniTalker, FLOAT, and Sonic. Despite their success, m

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction

DGX agent

arXiv:2605.25133v1 Announce Type: new Abstract: Reliably knowing when a language model is correct is almost as important as being correct. We introduce prover-verifier deliberation (PVD), an inference

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

DGX agent

arXiv:2509.25339v3 Announce Type: replace-cross Abstract: Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answerin

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

DGX agent

arXiv:2605.25981v1 Announce Type: new Abstract: We document an empirical phenomenon in chain-of-thought and ReAct agents driven by ten large language models from seven architecture families: meaning-b

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

DGX agent

arXiv:2605.25256v1 Announce Type: new Abstract: Aligning AI systems with organizational decision-making is typically framed as a single-target problem: make the model behave like the organization. We

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

World-State Transformations for Neuro-symbolic Interactive Storytelling

DGX agent

arXiv:2605.24719v1 Announce Type: cross Abstract: Large Language Models (LLMs) have changed the possibilities of Interactive Storytelling systems that process free-text user input. However, as more of

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering

DGX agent

arXiv:2605.23497v1 Announce Type: new Abstract: Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

Atom-level Protein Representation Learning Improves Protein Structure Prediction

DGX agent

arXiv:2605.22133v2 Announce Type: replace-cross Abstract: Recent advances in generative modeling show that pretrained representations can improve generation as conditioning features or alignment targe

model-releasesarxiv-cs-ai
25 May 2026
Safety

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training

DGX agent

arXiv:2603.06610v2 Announce Type: replace Abstract: Large language model (LLM) post-training enhances latent skills, unlocks value alignment, improves performance, and enables domain adaptation. Unfor

safetyarxiv-cs-lg
25 May 2026
Model Releases

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

DGX agent

arXiv:2605.23216v1 Announce Type: new Abstract: Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception t

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Convex Optimization for Alignment and Preference Learning on a Single GPU

DGX agent

arXiv:2605.23244v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) to align with human preferences has driven the success of systems such as Gemini and ChatGPT. However, approach

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

HARNESS-LM: A Three-Phase Training Recipe for Harnessing SLMs in Sponsored Search Retrieval

DGX agent

arXiv:2605.23572v1 Announce Type: cross Abstract: In the competitive landscape of sponsored search, balancing retrieval quality with production latency is a critical challenge. While large retrieval m

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness

DGX agent

arXiv:2605.23628v1 Announce Type: new Abstract: Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strate

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

LQ-rPPG: A Label-Quantized Coarse-to-Fine Learning Framework for Remote Physiological Measurement

DGX agent

arXiv:2605.23174v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) enables non-contact measurement of physiological signals from facial videos, offering strong potential for remote hea

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Parallel Context Compaction for Long-Horizon LLM Agent Serving

DGX agent

arXiv:2605.23296v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate growing conversation histories that eventually exceed the model's context window. Context compaction via LLM-based su

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks

DGX agent

arXiv:2605.23170v1 Announce Type: cross Abstract: Position-controlled evaluation is standard for retrieval tasks such as Needle-in-a-Haystack and RULER, but mainstream reasoning benchmarks do not cont

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

DGX agent

arXiv:2605.22903v1 Announce Type: cross Abstract: Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to wh

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

DGX agent

arXiv:2605.23904v1 Announce Type: new Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test

DGX agent

arXiv:2605.22841v1 Announce Type: cross Abstract: What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty c

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Understanding and Improving Noisy Embedding Techniques in Instruction Finetuning

DGX agent

arXiv:2605.23171v1 Announce Type: cross Abstract: Recent advancements in instructional fine-tuning have injected noise into embeddings, with NEFTune (Jain et al., 2024) setting benchmarks using unifor

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

DGX agent

arXiv:2605.22907v1 Announce Type: new Abstract: Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal s

model-releasesarxiv-cs-cv
25 May 2026
← Previous
1…320321322323324…1065
Next →