AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
86,965Total entries
1Added by human
86,964Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,106 results
Model Releases

Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens

DGX agent

arXiv:2602.13517v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated impressive reasoning capabilities by scaling test-time compute via long Chain-of-Thought (CoT). Howev

model-releasesarxiv-cs-cl
7 Jul 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

TRACE: Capability-Targeted Agentic Training

DGX agent

arXiv:2604.05336v2 Announce Type: replace Abstract: Models often fail to complete agentic tasks because they lack core capabilities required by the target environment. However, mainstream approaches f

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

DGX agent

arXiv:2511.20272v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of

model-releasesarxiv-cs-cv
7 Jul 2026
Safety

When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games

DGX agent

arXiv:2607.05132v1 Announce Type: cross Abstract: As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents tha

safetyarxiv-cs-cl
7 Jul 2026
Model Releases

When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

DGX agent

arXiv:2607.03836v1 Announce Type: cross Abstract: Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natu

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

DGX agent

arXiv:2607.03562v1 Announce Type: new Abstract: As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limi

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Framework

DGX agent

arXiv:2607.01581v1 Announce Type: new Abstract: The capacity of Large Language Models (LLMs) to reason about pedagogical intent within instructional communication remains underexplored, particularly i

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Black-Box Inference of LLM Architectural Properties with Restrictive API Access

DGX agent

arXiv:2607.01313v1 Announce Type: cross Abstract: In practice, most commercial LLM providers do not publicly release details of underlying LLM architectures. However, prior work has shown that given l

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

BuilderBench: The Building Blocks of Intelligent Agents

DGX agent

arXiv:2510.06288v4 Announce Type: replace Abstract: Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the limits set b

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

DecompRL: Solving Harder Problems by Learning Modular Code Generation

DGX agent

arXiv:2607.02390v1 Announce Type: new Abstract: How can Large Language Models (LLMs) solve problems they currently cannot? Repeated sampling scales test-time compute but GPU cost grows linearly with a

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Evidence-State Rewards for Long-Context Reasoning

DGX agent

arXiv:2607.02073v1 Announce Type: new Abstract: Long-context reasoning requires models to locate, revise, and synthesize evidence distributed across lengthy inputs. Existing long-context RL methods us

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

DGX agent

arXiv:2607.02396v1 Announce Type: new Abstract: Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behav

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories

DGX agent

arXiv:2607.02010v1 Announce Type: new Abstract: Multimodal large language models must adapt to evolving tasks and domains, yet continual improvement under bounded deployment footprint remains difficul

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding

DGX agent

arXiv:2601.01095v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand tempo

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

PACE: A Proxy for Agentic Capability Evaluation

DGX agent

arXiv:2607.02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation c

model-releasesarxiv-cs-ai
3 Jul 2026
Research

Self-Supervised Test-Time Tuning for Packet Loss Concealment

DGX agent

arXiv:2607.01823v1 Announce Type: cross Abstract: Packet loss concealment (PLC) reconstructs audio packets that are missing at the receiver, usually with a trained model whose parameters remain fixed

researcharxiv-cs-cl
3 Jul 2026
Hardware

WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs

DGX agent

arXiv:2607.02391v1 Announce Type: cross Abstract: Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requi

hardwarearxiv-cs-lg
3 Jul 2026
Model Releases

AutoMem: Automated Learning of Memory as a Cognitive Skill

DGX agent

arXiv:2607.01224v1 Announce Type: new Abstract: Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capacity known in cognitive science as m

model-releasesarxiv-cs-ai
2 Jul 2026
Safety

Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation

DGX agent

arXiv:2607.01208v1 Announce Type: cross Abstract: Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering user decisions at scale. Such pr

safetyarxiv-cs-ai
2 Jul 2026
Model Releases

DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for Autonomous Driving

DGX agent

arXiv:2607.00399v1 Announce Type: new Abstract: End-to-end autonomous driving models often encounter performance bottlenecks, as training-time scaling leads to high computational costs and diminishing

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Dual-Confidence Contrastive Decoding for Retrieval-Augmented Generation

DGX agent

arXiv:2607.00570v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) increasingly requires models to answer questions from multiple retrieved documents, where only some sources are rel

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

DGX agent

arXiv:2607.00218v1 Announce Type: cross Abstract: Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genu

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

FlowPath: Learning Data-Driven Manifolds with Invertible Flows for Robust Irregularly-sampled Time Series Classification

DGX agent

arXiv:2511.10841v3 Announce Type: replace-cross Abstract: Modeling continuous-time dynamics from sparse and irregularly-sampled time series remains a fundamental challenge. Neural controlled different

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine

DGX agent

arXiv:2607.00544v1 Announce Type: new Abstract: Reasoning segmentation requires localizing targets based on complex, implicit queries. Current end-to-end models typically entangle perception and deduc

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Linguistic Relative Policy Optimization for Video Anomaly Reasoning

DGX agent

arXiv:2607.00654v1 Announce Type: new Abstract: Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale an

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

DGX agent

arXiv:2607.01002v1 Announce Type: cross Abstract: In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pastin

model-releasesarxiv-cs-ai
2 Jul 2026
Research

Low Perplexity is Repetition: A One-Dimensional Self-Conditioning Attractor in Continuous Diffusion LMs

DGX agent

arXiv:2607.00588v1 Announce Type: new Abstract: Continuous diffusion language models such as ELF report record-low generative perplexity (Gen-PPL). We find a catch: these models repeat far more than h

researcharxiv-cs-cl
2 Jul 2026
Model Releases

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

DGX agent

arXiv:2510.24434v3 Announce Type: replace Abstract: The effectiveness of instruction-tuned Large Language Models (LLMs) is often limited in low-resource linguistic settings due to a lack of high-quali

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

Quantum vs. Classical Machine Learning: A Unified Empirical Comparison

DGX agent

arXiv:2607.01197v1 Announce Type: new Abstract: Quantum computing has emerged as a promising computational paradigm for machine learning (ML), with the potential to offer computational advantages over

model-releasesarxiv-cs-lg
2 Jul 2026
Safety

Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval

DGX agent

arXiv:2607.00448v1 Announce Type: cross Abstract: The two-tower model has been widely used for large-scale recommendation systems, particularly in the retrieval stage. Industry standards for training

safetyarxiv-cs-ai
2 Jul 2026
Model Releases

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration

DGX agent

arXiv:2607.00816v1 Announce Type: new Abstract: High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when t

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

DGX agent

arXiv:2511.17731v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

WorkBench Revisited: Workplace Agents Two Years On

DGX agent

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

DGX agent

arXiv:2607.00664v1 Announce Type: new Abstract: We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese

model-releasesarxiv-cs-cl
2 Jul 2026
Safety

Addressing Over-Refusal in LLMs with Competing Rewards

DGX agent

arXiv:2606.31748v1 Announce Type: new Abstract: Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones. Tho

safetyarxiv-cs-lg
1 Jul 2026
Model Releases

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

DGX agent

arXiv:2606.31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) model

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Efficient Public Verification of Private ML via Regularization

DGX agent

arXiv:2512.04008v2 Announce Type: replace Abstract: Training with differential privacy (DP) guarantees dataset members that they cannot be identified by users of the released model. However, those dat

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents

DGX agent

arXiv:2606.31270v1 Announce Type: cross Abstract: Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant atten

model-releasesarxiv-cs-ai
1 Jul 2026
Research

MAPE: Defending Against Transferable Adversarial Attacks Using Multi-Source Adversarial Perturbations Elimination

DGX agent

arXiv:2606.31378v1 Announce Type: new Abstract: Neural networks are vulnerable to meticulously crafted adversarial examples, leading to high-confidence misclassifications in image classification tasks

researcharxiv-cs-cv
1 Jul 2026
Model Releases

Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2

DGX agent

arXiv:2606.31543v1 Announce Type: new Abstract: Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles

DGX agent

arXiv:2606.23672v2 Announce Type: replace Abstract: This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In this tas

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Temperature Field Reconstruction of Tungsten Monoblock Divertor on EAST using Physics-aware Neural Operator Transformer

DGX agent

arXiv:2606.31574v1 Announce Type: cross Abstract: Accurate modeling of the divertor temperature field is essential for preventing material melting and damage and for extending the service life of fusi

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation

DGX agent

arXiv:2606.31704v1 Announce Type: new Abstract: The deployment of face detection models in real-world applications raises important fairness concerns, as these systems may showcase performance dispari

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

A Machine-Verified Proof of a Quantum-Optimization Conjecture

DGX agent

arXiv:2606.29687v1 Announce Type: cross Abstract: We report a machine-verified resolution of a problem open for over a decade in quantum optimization: the Farhi, Goldstone and Gutmann (FGG) conjecture

model-releasesarxiv-cs-ai
30 Jun 2026
Hardware

A Trainable-by-Parts Operator Learning Framework: Bridging DeepONet and Karhunen-Loeve Expansions for Large-Scale Applications

DGX agent

arXiv:2606.28519v1 Announce Type: new Abstract: Training operator-learning models for large-scale problems governed by partial differential equations (PDEs) is challenging due to the curse of dimensio

hardwarearxiv-cs-lg
30 Jun 2026
Model Releases

Adaptive Financial Transformer with Regime-Gated Attention for Stock Return Prediction

DGX agent

arXiv:2606.29347v1 Announce Type: cross Abstract: Adaptive Financial Transformer (AFT) is proposed for stock return prediction under non-stationary financial markets. The model incorporates a Market R

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

Agent Safety Is Action Alignment

DGX agent

arXiv:2606.28739v1 Announce Type: new Abstract: Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe,

safetyarxiv-cs-ai
30 Jun 2026
Research

ArchesClimate: Probabilistic Decadal Ensemble Generation With Flow Matching

DGX agent

arXiv:2509.15942v3 Announce Type: replace-cross Abstract: Internal variability is a dominant contributor to the uncertainty of predictions at the interannual to decadal timescale. A typical approach t

researcharxiv-cs-ai
30 Jun 2026
← Previous
1…311312313314315…1065
Next →