AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,566 results
Model Releases

The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

DGX agent

arXiv:2605.08427v1 Announce Type: new Abstract: Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in

model-releasesarxiv-cs-ai
12 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

The autoPET3 Challenge: Automated Lesion Segmentation in Whole-Body PET/CT nicode{x2013} Multitracer Multicenter Generalization

DGX agent

arXiv:2605.05775v2 Announce Type: replace-cross Abstract: We report the design and results of the third autoPET challenge (MICCAI 2024), which benchmarked automated lesion segmentation in whole-body P

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Differences Between Direct Alignment Algorithms are a Blur

DGX agent

arXiv:2502.01237v3 Announce Type: replace Abstract: Direct Alignment Algorithms (DAAs) simplify LLM alignment by directly optimizing policies, bypassing reward modeling and RL. While DAAs differ in th

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

The Echo Amplifies the Knowledge: Somatic Marker Analogues in Language Models via Emotion Vector Re-Injection

DGX agent

arXiv:2605.08611v1 Announce Type: new Abstract: Current language model memory systems store what happened but not how it felt. This distinction -- between semantic memory (knowing about a past event)

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs

DGX agent

arXiv:2605.08737v1 Announce Type: cross Abstract: On-policy distillation (OPD) is widely used for LLM post-training. When pushed with a reward-extrapolation coefficient lambda > 1, the student can lif

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws

DGX agent

arXiv:2605.09887v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) operationalise the linear representation hypothesis: they reconstruct model activations as sparse linear combinations of in

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations

DGX agent

arXiv:2605.09195v1 Announce Type: new Abstract: Large language models confidently produce outdated answers, and no existing method can detect them. We show this is not an engineering failure but a str

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning

DGX agent

arXiv:2605.08746v1 Announce Type: new Abstract: In training a neural network with gradient descent (GD), each iteration induces a linear operator that governs first-order updates to a model's internal

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark

DGX agent

arXiv:2605.09900v1 Announce Type: new Abstract: A vision-language model can look at a knot diagram and report what it sees, yet fail to act on that structure. KnotBench pairs an 858,318-image corpus f

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies

DGX agent

arXiv:2605.10799v1 Announce Type: cross Abstract: Corruption studies, the primary tool for evaluating chain-of-thought (CoT) faithfulness, identify which chain positions are 'computationally important

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs

DGX agent

arXiv:2605.09844v1 Announce Type: new Abstract: The Metacognitive Probe is an exploratory five-task, 15-slot diagnostic that decomposes an LLM's confidence behaviour into five behaviourally-distinct d

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Realignment Problem: When Right becomes Wrong in LLMs

DGX agent

arXiv:2511.02623v2 Announce Type: replace Abstract: Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over tim

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

The silent removal of Study Mode from ChatGPT is a big mistake (both Claude and Gemini still have theirs) We have enough evidence that using…

DGX agent

The silent removal of Study Mode from ChatGPT is a big mistake (both Claude and Gemini still have theirs) We have enough evidence that using AI in assistant mode to study can hurt learning because it

model-releasesethan-mollick--x
12 May 2026
Model Releases

The Silent Vote: Improving Zero-Shot LLM Reliability by Aggregating Semantic Neighborhoods

DGX agent

arXiv:2605.09739v1 Announce Type: cross Abstract: Large Language Models are increasingly used as zero-shot classifiers in complex reasoning tasks. However, standard constrained decoding suffers from a

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory

DGX agent

arXiv:2605.09330v1 Announce Type: cross Abstract: Agentic memory enables LLMs to persist information beyond a single context window and reuse it in later decisions, but it also introduces a new vulner

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The US House Oversight Committee launches a probe into potential conflicts in Sam Altman's personal investments; letter: several GOP AGs call for an SEC review (Wall Street Journal)

DGX agent

Wall Street Journal: The US House Oversight Committee launches a probe into potential conflicts in Sam Altman's personal investments; letter: several GOP AGs call for an SEC review — Republican-led Ho

model-releasestechmeme
12 May 2026
Model Releases

The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models

DGX agent

arXiv:2601.02954v3 Announce Type: replace-cross Abstract: Large audio-language models have made rapid progress in recognizing what is present in an audio clip, but spatial audio-language understanding

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Wristband Gaussian Loss: Deterministic, Composable Latents via a Sphere-Interval Decomposition

DGX agent

arXiv:2605.08749v1 Announce Type: new Abstract: We present the Wristband Gaussian Loss, a deterministic batch loss for Gaussianizing point embeddings without sampling, KL terms, or iterative transport

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization,…

DGX agent

This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization, custom kernels, and rack-scale NVLink turn GB200 into faste

model-releasesperplexity--x
12 May 2026
Model Releases

ThreatCore: A Benchmark for Explicit and Implicit Threat Detection

DGX agent

arXiv:2605.10563v1 Announce Type: cross Abstract: Threat detection in Natural Language Processing lacks consistent definitions and standardized benchmarks, and is often conflated with broader phenomen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning

DGX agent

arXiv:2605.09544v1 Announce Type: new Abstract: Tool-integrated reasoning has emerged as a promising paradigm for enhancing large language models with external computation, retrieval, and execution ca

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TIDES: Implicit Time-Awareness in Selective State Space Models

DGX agent

arXiv:2605.09742v1 Announce Type: cross Abstract: Selective state space models (SSMs), such as Mamba, achieve strong per-token expressivity by making the time discretization step Tilde{Delta} a learne

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Tight Generalization Bounds for Noiseless Inverse Optimization

DGX agent

arXiv:2605.08866v1 Announce Type: cross Abstract: Inverse optimization (IO) seeks to infer the parameters of a decision-maker's objective from observed context--action data. We study noiseless IO, whe

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

TiledAttention: a CUDA Tile SDPA Kernel for PyTorch

DGX agent

arXiv:2603.01960v2 Announce Type: replace-cross Abstract: TiledAttention is a scaled dot-product attention (SDPA) forward operator for SDPA research on NVIDIA GPUs. Implemented in cuTile Python (TileI

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TINS: Test-time ID-prototype-separated Negative Semantics Learning for OOD Detection

DGX agent

arXiv:2605.10756v1 Announce Type: new Abstract: Vision-language models enable OOD detection by comparing image alignment with ID labels and negative semantics. Existing negative-label-based methods ma

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

To Redact, or not to Redact? A Local LLM Approach to Deliberative Process Privilege Classification

DGX agent

arXiv:2605.10211v1 Announce Type: cross Abstract: Government transparency laws, like the Freedom of Information (FOIA) acts in the United States and United Kingdom, and the Woo (Open Government Act) i

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models

DGX agent

arXiv:2605.09904v1 Announce Type: new Abstract: Video large language models (Video-LLMs) have achieved remarkable progress in general video understanding, yet their ability to maintain temporal object

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation

DGX agent

arXiv:2605.08541v1 Announce Type: new Abstract: Neural scaling laws approximate a language model's loss as a power-law function of parameter count N and token count D. Following Chinchilla-style compu

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Towards a Large Language-Vision Question Answering Model for MSTAR Automatic Target Recognition

DGX agent

arXiv:2605.10772v1 Announce Type: cross Abstract: Large language-vision models (LLVM), such as OpenAI's ChatGPT and GPT-4, have gained prominence as powerful tools for analyzing text and imagery. The

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Towards Compact Sign Language Translation: Frame Rate and Model Size Trade-offs

DGX agent

arXiv:2605.09554v1 Announce Type: new Abstract: Sign Language Translation (SLT) converts sign language videos into spoken-language text, bridging communication between Deaf and hearing communities. Cu

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Towards Conversational Medical AI with Eyes, Ears and a Voice

DGX agent

arXiv:2605.09272v1 Announce Type: new Abstract: The practice of medicine relies not only upon skillful dialogue but also on the nuanced exchange and interpretation of rich auditory and visual cues bet

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Towards Cross-lingual Values Judgment: A Consensus-Pluralism Perspective

DGX agent

arXiv:2602.17283v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are employed worldwide, existing evaluation paradigms for their multilingual capabilities primarily focus on f

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse

DGX agent

arXiv:2605.09965v1 Announce Type: new Abstract: The real world unfolds along a single set of physics laws, yet human intelligence demonstrates a remarkable capacity to generalize experiences from this

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Towards Generative Predictive Display for Vision-Based Teleoperation: A Zero-Shot Benchmark of Off-the-Shelf Video Models

DGX agent

arXiv:2605.09670v1 Announce Type: cross Abstract: Teleoperation systems are fundamentally limited by communication latency, which degrades situational awareness and control performance. Predictive dis

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

DGX agent

arXiv:2605.10832v1 Announce Type: new Abstract: Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visua

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs

DGX agent

arXiv:2604.17502v2 Announce Type: replace Abstract: Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length trajectori

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias

DGX agent

arXiv:2605.09087v1 Announce Type: cross Abstract: Audio deepfake detection systems are increasingly deployed in high-stakes security applications, yet their fairness across demographic groups remains

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Towards Universal Gene Regulatory Network Inference: Unlocking Generalizable Regulatory Knowledge in Single-cell Foundation Models

DGX agent

arXiv:2605.08128v1 Announce Type: cross Abstract: Gene Regulatory Network (GRN) inference is essential for understanding complex cellular mechanisms, rendered tractable through single-cell transcripto

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

DGX agent

arXiv:2605.09934v1 Announce Type: new Abstract: Multimodal large language models increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, a

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Tracing Moral Foundations in Large Language Models

DGX agent

arXiv:2601.05437v2 Announce Type: replace-cross Abstract: Large language models often produce human-like moral judgments, but it is unclear whether this reflects an internal conceptual structure or su

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models

DGX agent

arXiv:2605.08974v1 Announce Type: cross Abstract: While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We arg

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Trajectory Supervision for Continual Tool-Use Learning in LLMs

DGX agent

arXiv:2605.09734v1 Announce Type: cross Abstract: Most language-model training data shows final artifacts, not the process that produced them. We study a tractable version of this question in tool use

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TrajPrism: A Multi-Task Benchmark for Language-Grounded Urban Trajectory Understanding

DGX agent

arXiv:2605.10782v1 Announce Type: new Abstract: Urban mobility is naturally expressed both as trajectories in space and as natural-language descriptions of travel intent, constraints, and preferences.

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training

DGX agent

arXiv:2605.10835v1 Announce Type: new Abstract: Optical Music Recognition (OMR), the task of transcribing sheet music into a structured textual representation, is currently bottlenecked by a lack of l

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Transfer Learning of Multiobjective Indirect Low-Thrust Trajectories Using Diffusion Models and Markov Chain Monte Carlo

DGX agent

arXiv:2605.09125v1 Announce Type: cross Abstract: Preliminary low-thrust spacecraft mission design is a global search problem characterized by a complex solution landscape, multiple objectives, and nu

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models

DGX agent

arXiv:2601.22478v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has become the dominant method for reinforcement learning with verifiable rewards in large language models

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Transformers Provably Learn Sparse XOR with Polylogarithmic Parameters

DGX agent

arXiv:2502.07553v2 Announce Type: replace Abstract: Learning sparse parity functions has become a theoretical testbed for studying feature learning in neural networks. However, existing analyses prima

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Try Grok Voice

DGX agent

Try Grok Voice Grok Voice Think Fast 1.0 ranks #1 on the Artificial Analysis τ-Voice benchmark for real-world agentic customer service resolution Absolutely outperforming GPT-Realtime-2 (High) and Gem

model-releaseselon-musk--x
12 May 2026
← Previous
1…335336337338339…471
Next →