AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
29 Jun 2026

Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems

ApplicationsDGX agent

arXiv:2606.27797v1 Announce Type: cross Abstract: Knowledge Distillation (KD) enables training smaller student models under the guidance of larger teacher models, and the widely adopted TRL library im

Speculative Refinement: A Hybrid Autoregressive Diffusion Decoding Strategy and Its Behavior Across Benchmarks

Model ReleasesDGX agent

arXiv:2606.27474v1 Announce Type: cross Abstract: How should we evaluate generation systems that combine autoregressive (AR) and diffusion decoding? We study this question through Speculative Refineme

27 Jun 2026

How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and cach…

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
TutorialsDGX agent

How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching. Better Defaults (not Usage Caps) – Engineers can choose

26 Jun 2026

Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare

Model ReleasesDGX agent

arXiv:2606.26104v1 Announce Type: cross Abstract: Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about ani

Autoregressive Boltzmann Generators

Model ReleasesDGX agent

arXiv:2606.27361v1 Announce Type: cross Abstract: Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics. This challenge has driven the dev

Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT

Model ReleasesDGX agent

arXiv:2606.26861v1 Announce Type: new Abstract: Deploying large language models (LLMs) on Industrial Internet of Things (IIoT) edge devices demands extreme compression, yet existing structured pruning

Context Recycling for Long-Horizon LLM Inference

Model ReleasesDGX agent

arXiv:2606.26105v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational horizons due t

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Model ReleasesDGX agent

arXiv:2606.27264v1 Announce Type: new Abstract: Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text jud

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues

Model ReleasesDGX agent

arXiv:2606.26602v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive fine-grained perception capabilities. However, existing ben

Empirical Software Engineering TerraProbe: A Layered-Oracle Framework for Detecting Deceptive Fixes in LLM-Assisted Terraform

Model ReleasesDGX agent

arXiv:2606.26590v1 Announce Type: new Abstract: Security misconfigurations in Terraform Infrastructure-as-Code are a growing risk in cloud deployments, and large language models are increasingly used

From Guessing to Placeholding: A Cost-Theoretic Framework for Uncertainty-Aware Code Completion

Model ReleasesDGX agent

arXiv:2604.01849v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated exceptional proficiency in code completion, they typically adhere to a Hard Completion (HC) par

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training

Model ReleasesDGX agent

arXiv:2606.26102v1 Announce Type: cross Abstract: Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these process

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

Model ReleasesDGX agent

arXiv:2606.26346v1 Announce Type: new Abstract: Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-doma

Implementation of reinforcement learning in chemical reaction networks: application to phototaxis as curiosity-driven exploration

Model ReleasesDGX agent

arXiv:2606.26168v1 Announce Type: new Abstract: Living systems navigate environments using noisy and incomplete sensory signals. In unicellular algae, phototaxis is often modeled as a mechanistic run-

Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

Model ReleasesDGX agent

arXiv:2606.26498v1 Announce Type: cross Abstract: This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unkn

OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference

Model ReleasesDGX agent

arXiv:2601.13300v2 Announce Type: replace Abstract: Benchmarking large language models (LLMs) is critical for understanding their capabilities, limitations, and robustness. In addition to interface ar

OpenAI introduces GPT-5.6 to challenge Claude Mythos 5

Model ReleasesDGX agent

OpenAI Group PBC today introduced GPT-5.6, a new series of large language models that it says can outperform Claude Mythos 5 across certain coding tasks. The most advanced algorithm in the lineup is k

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

Model ReleasesDGX agent

arXiv:2606.26552v1 Announce Type: cross Abstract: The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread

PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

Model ReleasesDGX agent

arXiv:2606.26551v1 Announce Type: new Abstract: While instruction-based image editing, enabled by multi-modal generative models, has advanced significantly, existing benchmarks lack a comprehensive ev

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

Model ReleasesDGX agent

arXiv:2601.11061v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is highly effective for enhancing LLM reasoning, yet recent evidence shows models like Q

What We are Missing in Multimodal LLM Evaluation?

Model ReleasesDGX agent

arXiv:2606.26348v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their c

25 Jun 2026

2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction

Model ReleasesDGX agent

arXiv:2603.19964v3 Announce Type: replace Abstract: High-resolution geometric prediction is essential for robust perception in autonomous driving, robotics, and AR/MR, but current foundation models ar

BOFA: Bridge-Layer Orthogonal Low-Rank Fusion for CLIP-Based Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2511.11421v2 Announce Type: replace Abstract: Class-Incremental Learning (CIL) aims to continually learn new categories without forgetting previously acquired knowledge. Vision-language models s

C3-Bench: A Context-Aware Change Captioning Benchmark

Model ReleasesDGX agent

arXiv:2606.25445v1 Announce Type: new Abstract: While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world chang

CausalRAG2: Hierarchical Causal Knowledge Graph Design for RAG

Model ReleasesDGX agent

arXiv:2602.05143v2 Announce Type: replace Abstract: Retrieval augmented generation (RAG) has enhanced large language models by enabling access to external knowledge, with graph-based RAG emerging as a

Curvature-Guided Mixing for MLLM Adaptation

Model ReleasesDGX agent

arXiv:2606.24963v1 Announce Type: new Abstract: Fine-tuning Multimodal Large Language Models (MLLMs) on specialized tasks often leads to catastrophic forgetting of their general capabilities. Existing

Dream at SemEval-2026 Task 13: SALSA for Single-Pass Machine-Generated Code Detection

Model ReleasesDGX agent

arXiv:2606.25102v1 Announce Type: new Abstract: Large language models have transformed code generation, raising concerns around authorship, assessment integrity, and software trust. SemEval-2026 Task

Evidence for feature-specific error correction in LLMs

Model ReleasesDGX agent

arXiv:2606.24964v1 Announce Type: new Abstract: Understanding the features of large language models (LLMs) is a central goal of interpretability. LLMs are commonly assumed to use superposition to repr

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice

Model ReleasesDGX agent

arXiv:2606.22327v2 Announce Type: replace Abstract: The explosive demand for interactive Large Language Model serving has highlighted the management of the Key-Value cache's dynamic memory footprint a

Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation

ResearchDGX agent

arXiv:2606.25278v1 Announce Type: new Abstract: Multi-modal fusion and multi-model ensembling are prevalent in enhancing the performance of 3D semantic segmentation. Despite the impressive performance

LLM Performance on a Real, Double-Marked GCSE Benchmark

Model ReleasesDGX agent

arXiv:2606.24973v1 Announce Type: new Abstract: We introduce a dataset of 32,534 double-marked real student responses to GCSE mock exams (GCSEs are the UK's national exams, taken at age ~16), spanning

Operator Boosting Produces Pareto-Efficient PDE Surrogates

Model ReleasesDGX agent

arXiv:2606.17460v2 Announce Type: replace Abstract: Neural operators are widely used as surrogate solution maps for partial differential equations (PDEs), but full-size models can be costly to store,

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

Model ReleasesDGX agent

arXiv:2606.18936v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature a

Transferable Attack against Face Swapping in an Extended Space

ResearchDGX agent

arXiv:2606.25376v1 Announce Type: new Abstract: Although deep Face Swapping (FS) models may benefit the entertainment industry, they pose severe threats to privacy and security. Existing protections,

Verifiable Manifest Signing and Transparency Enforcement for Secure MCP-Based LLM Pipelines

Model ReleasesDGX agent

arXiv:2601.23132v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed in tool-driven environments such as healthcare analytics, financial systems, retrieval-

24 Jun 2026

A Benchmark for Hallucination Detection in VLMs for Gastrointestinal Endoscopy

Model ReleasesDGX agent

arXiv:2606.24115v1 Announce Type: cross Abstract: Vision-language models (VLMs) are prone to hallucination, which remains a major barrier to their safe deployment in clinical practice. To date, most h

AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning

Model ReleasesDGX agent

arXiv:2606.24526v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grou

BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming

Model ReleasesDGX agent

arXiv:2606.24740v1 Announce Type: new Abstract: Recent advances in vision-language models (VLMs) such as CLIP have demonstrated strong generalization across natural-image domains. However, adapting th

Dual-Branch Cross-Projection Debiasing through Diffusion-based Disentanglement

Model ReleasesDGX agent

arXiv:2606.24161v1 Announce Type: new Abstract: Foundation models trained on biased datasets often rely on spurious correlations between target labels and non-causal attributes, resulting in poor gene

Flow-Corrected Thompson Sampling for Non-Stationary Contextual Bandits

Model ReleasesDGX agent

arXiv:2606.23933v1 Announce Type: cross Abstract: We study non-stationary linear contextual bandits where the reward model drifts over time, rendering classical contextual bandit algorithms brittle be

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.24133v1 Announce Type: cross Abstract: The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of Large Language Model (LLM) pre-t

Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching

ApplicationsDGX agent

arXiv:2606.24457v1 Announce Type: new Abstract: Recent advances in stereo matching have achieved remarkable accuracy, but often rely on large models, heavy computation, or additional foundation-model

Quantum ring all-reduce: communication and privacy advantages for distributed learning

Model ReleasesDGX agent

arXiv:2606.20344v2 Announce Type: replace-cross Abstract: Machine learning models have scaled to unprecedented sizes, making training across distributed devices the de facto standard in the field. In

Reasoning as Attractor Dynamics: Latent Memory Retrieval via Gibbs-Weighted Energy Minimization

Model ReleasesDGX agent

arXiv:2606.24543v1 Announce Type: new Abstract: Large Language Models (LLMs) are traditionally viewed as autoregressive generators. However, from the perspective of collective computation, they functi

Towards Spec Learning: Inference-Time Alignment from Preference Pairs

Model ReleasesDGX agent

arXiv:2606.24004v1 Announce Type: cross Abstract: Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on a careful

When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs

Model ReleasesDGX agent

arXiv:2606.24370v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into decision-support roles in business and policy contexts. While prior benchmark studies have

ZONOS2 Technical Report

Model ReleasesDGX agent

arXiv:2606.24320v1 Announce Type: cross Abstract: We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0

23 Jun 2026

A Verifiable Search Is Not a Learnable Chain-of-Thought

Model ReleasesDGX agent

arXiv:2606.21884v1 Announce Type: new Abstract: It is tempting to assume any task solvable by a short program can be taught to a model as its chain-of-thought: write the steps out, fine-tune, and the

ASCII Art Turns LLMs into VLA Controllers

Model ReleasesDGX agent

arXiv:2606.21470v1 Announce Type: cross Abstract: Vision--Language--Action (VLA) controllers are often built by extending vision--language models (VLMs) with action supervision, relying on multimodal

Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies

Model ReleasesDGX agent

arXiv:2606.20599v1 Announce Type: cross Abstract: Tree of Thought (ToT) search has become a promising direction for improving the reasoning capabilities of large language models, but deploying these m

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR

Model ReleasesDGX agent

arXiv:2606.20770v1 Announce Type: cross Abstract: Current Vision-Language Models (VLMs) are celebrated for their multilingual capabilities, yet they operate under a flawed assumption: that one languag

Chehre: An Emoji-Prompted Video Dataset for Perceptually Diverse Facial Expression Recognition

Model ReleasesDGX agent

arXiv:2606.21657v1 Announce Type: new Abstract: Facial expressions are nonverbal social signals used in human interaction, but facial expression recognition datasets often focus on static images, basi

Chem2Gen-Bench: Benchmarking Chemical-to-Genetic Translation in Perturbation Response Space

Model ReleasesDGX agent

arXiv:2606.21109v1 Announce Type: new Abstract: Virtual-cell and perturbation models are increasingly used to predict cellular responses for biomedical discovery, but chemical and genetic perturbation

Confidently Wrong: Severity-Aware Calibration of Prompt-Injection Detectors under Attack Shift

Model ReleasesDGX agent

arXiv:2606.22659v1 Announce Type: cross Abstract: Prompt-injection detectors are deployed as guards: a model scores an input and a downstream system trusts or blocks it on that score. I study the conf

Double-Diffusion: Balancing Speed, Accuracy, and Uncertainty in Probabilistic Forecasting for Urban Sensor Networks

Model ReleasesDGX agent

arXiv:2506.23053v3 Announce Type: replace Abstract: Urban sensor networks need forecasts that are accurate, carry useful uncertainty, and refresh fast enough to act on as new readings arrive. These go

ELDiff: When Evidential Learning Meets Text-to-Image Diffusion

Model ReleasesDGX agent

arXiv:2606.20924v1 Announce Type: new Abstract: In multi-object text-to-image (T2I) diffusion, ensuring semantic consistency between textual prompts and generated visual content is crucial for image s

Fara-1.5: Scalable Learning Environments for Computer Use Agents

Model ReleasesDGX agent

arXiv:2606.20785v1 Announce Type: cross Abstract: Collecting computer use data from human demonstrations is expensive and slow, motivating the need for scalable generation strategies. This requires tw

Good-Enough LLM Obfuscation (GELO)

Model ReleasesDGX agent

arXiv:2603.05035v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly served on shared accelerators where an adversary with read access to device memory can observe K

Happy Young Women, Grumpy Old Men? Emotion-Driven Demographic Biases in Synthetic Face Generation

SafetyDGX agent

arXiv:2602.00032v3 Announce Type: replace-cross Abstract: Synthetic faces from text-to-image (T2I) models pervade digital media, yet their demographic biases under emotionally conditioned prompts rema

Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping

TutorialsDGX agent

arXiv:2606.23682v1 Announce Type: new Abstract: Reference-based diffusion models enable highly controllable image generation by leveraging elements from input images to guide prompt-driven synthesis.

← Previous
1…303304305306307…1036
Next →