AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
49,435 results
Model Releases

SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

DGX agent

arXiv:2608.12127v1 Announce Type: new Abstract: Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is

model-releasesarxiv-cs-cv
13 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

DGX agent

arXiv:2608.11694v1 Announce Type: cross Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same pr

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates

DGX agent

arXiv:2608.11435v1 Announce Type: new Abstract: Forward and inverse modeling of parametric dynamical systems requires surrogate models that are not only accurate for state prediction, but also informa

model-releasesarxiv-cs-lg
13 Aug 2026
Agents

Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning

DGX agent

arXiv:2608.10438v1 Announce Type: new Abstract: Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. Fo

agentsarxiv-cs-ai
12 Aug 2026
Model Releases

DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

DGX agent

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

DGX agent

arXiv:2608.10764v1 Announce Type: new Abstract: Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

DGX agent

arXiv:2608.10300v1 Announce Type: new Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human re

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator

DGX agent

arXiv:2608.09971v1 Announce Type: cross Abstract: Over the past few years, the rapid development of machine learning (ML) models for weather forecasting has produced deterministic models whose medium-

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

DGX agent

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent

DGX agent

arXiv:2608.10258v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

DGX agent

arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached o

model-releasesarxiv-cs-lg
12 Aug 2026
Model Releases

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

DGX agent

arXiv:2608.09164v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms.

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

FEAST: Federated Shared-Space Training for Resource-Heterogeneous Clients

DGX agent

arXiv:2608.09250v1 Announce Type: new Abstract: Federated learning (FL) must serve devices with varying computational capabilities. A fixed model cannot suit all devices, while training one model per

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Staying True to the Origin: Continuous Image Stylization with Smooth Transitions

DGX agent

arXiv:2608.08125v1 Announce Type: new Abstract: Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a g

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

DGX agent

arXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the rig

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

DGX agent

arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation

DGX agent

arXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

DGX agent

arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, c

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Explanation Stability of Test-Time Adaptation in Computational Pathology: A Large-Scale Benchmark

DGX agent

arXiv:2608.07062v1 Announce Type: new Abstract: Test-time adaptation (TTA) has become a practical way to adapt deployed models to unlabeled target data, a setting that is especially relevant in comput

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

Generative Embedding Benchmark: How Much Information Survives in a Dense Embedding?

DGX agent

arXiv:2608.06972v1 Announce Type: new Abstract: Embeddings have emerged as a standard representational interface linking foundation models with downstream systems. Most embedding benchmarks assess rep

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

The Bitter Lesson of Tool Calling

DGX agent

arXiv:2608.06370v1 Announce Type: new Abstract: Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by

model-releasesarxiv-cs-cl
7 Aug 2026
Model Releases

Item Response Theory for AI Safety

DGX agent

arXiv:2608.05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to tr

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

DGX agent

arXiv:2608.03892v1 Announce Type: new Abstract: We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, r

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

SocietyBench: Forecasting Counterfactual Social-World Evolution

DGX agent

arXiv:2608.04009v1 Announce Type: new Abstract: Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a b

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

DGX agent

arXiv:2512.08163v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are increasingly used as functional models of human vision, yet standard monocular depth estimation (MDE) benchmarks lar

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Evaluating VLMs on Multimodal Aristotelian Persuasion Tasks

DGX agent

arXiv:2608.01238v1 Announce Type: new Abstract: Vision Language Models (VLMs) have demonstrated exceptional performance across various tasks. However, they have not yet been thoroughly evaluated on mo

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Image-Space Rule Discovery

DGX agent

arXiv:2608.00490v1 Announce Type: new Abstract: Can image-editing models discover visual rules in image space and complete problem-solving end-to-end? We tackle this question in the spirit of a human

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

DGX agent

arXiv:2608.01522v1 Announce Type: cross Abstract: Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning t

researcharxiv-cs-cl
4 Aug 2026
Model Releases

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

DGX agent

arXiv:2607.29549v1 Announce Type: new Abstract: Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challen

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

DGX agent

arXiv:2607.28658v1 Announce Type: cross Abstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, e

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction

DGX agent

arXiv:2607.28079v1 Announce Type: new Abstract: Chemical property prediction plays a critical role in accelerating scientific discovery in chemistry, materials science, and drug development. However,

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

DGX agent

arXiv:2607.27919v1 Announce Type: new Abstract: Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independent

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

DGX agent

arXiv:2607.27378v1 Announce Type: new Abstract: Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-g

model-releasesarxiv-cs-cv
31 Jul 2026
Model Releases

APEX-Accounting

DGX agent

arXiv:2607.27189v1 Announce Type: new Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountant

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

GPT-Red: Automated Red Teaming via Self-Play at Scale

DGX agent

arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce extbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

DGX agent

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation

DGX agent

arXiv:2509.22768v3 Announce Type: replace Abstract: We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach

DGX agent

arXiv:2607.26608v1 Announce Type: new Abstract: Training-free fusion of heterogeneous multimodal large language models (MLLMs) provides a direct route for cross-scale capability transfer, yet improvem

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization

DGX agent

arXiv:2607.25451v1 Announce Type: new Abstract: Language models are almost always quantized before they are deployed, and a growing line of work asks whether quantization also lowers their privacy ris

model-releasesarxiv-cs-lg
29 Jul 2026
Model Releases

Emergent Latent-State Computation under Stochastic Volatility

DGX agent

arXiv:2607.25459v1 Announce Type: cross Abstract: Mechanistic interpretability has largely focused on language models and deterministic toy tasks. Much less is known about how sequence models internal

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising

DGX agent

arXiv:2607.24779v1 Announce Type: new Abstract: Online advertising bidding systems typically deploy multiple offline-trained expert models (e.g., PID controllers, model predictive control, offline RL

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

Auditing Alignment Controllability in LLMs via Political Axes

DGX agent

arXiv:2607.23519v1 Announce Type: cross Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in dep

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

DGX agent

arXiv:2607.23159v1 Announce Type: new Abstract: Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although most are discard

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

DGX agent

arXiv:2607.23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

DGX agent

arXiv:2607.23802v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-sc

model-releasesarxiv-cs-ai
28 Jul 2026
Local Ai

Generative Video Compression with Adaptive Score Distillation

DGX agent

arXiv:2607.22772v1 Announce Type: cross Abstract: Diffusion models provide strong generative capabilities for video compression at ultra-low bitrates. Existing diffusion-based video codecs adapt base

local-aiarxiv-cs-cv
28 Jul 2026
Model Releases

LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning

DGX agent

arXiv:2607.22777v1 Announce Type: cross Abstract: Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along amino-acid se

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion

DGX agent

arXiv:2607.22696v1 Announce Type: cross Abstract: High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget of a singl

model-releasesarxiv-cs-ai
28 Jul 2026
← Previous
1…196197198199200…1030
Next →