AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
12 May 2026

Sinkhorn Treatment Effects: A Causal Optimal Transport Measure

Model ReleasesDGX agent

arXiv:2605.08485v1 Announce Type: cross Abstract: We introduce the Sinkhorn treatment effect, an entropic optimal transport measure of divergence between counterfactual distributions. Unlike classical

Sketch-and-Verify: Structured Inference-Time Scaling via Program Sketching

Model ReleasesDGX agent

arXiv:2605.08658v1 Announce Type: cross Abstract: SKETCHVERIFY is a within-tier cost-performance policy, not a universal accuracy improvement. The operational question: a practitioner stuck with a sma

SLAM: Structural Linguistic Activation Marking for Language Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.05443v2 Announce Type: replace-cross Abstract: LLM watermarks must be detectable without compromising text quality, yet most existing schemes bias the next-token distribution and pay for de

SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2605.10376v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliabl

SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing

Model ReleasesDGX agent

arXiv:2605.10831v1 Announce Type: cross Abstract: Large language models possess strong chemical reasoning capabilities, making them effective molecular editors. However, property-relevant information

SmartEval: A Benchmark for Evaluating LLM-Generated Smart Contracts from Natural Language Specifications

Model ReleasesDGX agent

arXiv:2605.09610v1 Announce Type: cross Abstract: We introduce SmartEval, a benchmark for systematically evaluating the quality of Solidity smart contracts generated by large language models (LLMs) fr

SMIXAE: Towards Unsupervised Manifold Discovery in Language Models

Model ReleasesDGX agent

arXiv:2605.09224v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have been used widely to decompose and interpret neural network activations, especially those of transformer language models.

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

Model ReleasesDGX agent

arXiv:2605.09598v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos du

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Model ReleasesDGX agent

arXiv:2605.09063v1 Announce Type: new Abstract: Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challengi

Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection

Model ReleasesDGX agent

arXiv:2605.08583v1 Announce Type: new Abstract: Large language models are increasingly used in scientific writing, yet they can fabricate citation-shaped references that appear plausible but fail bibl

Sources: Anthropic is in talks to raise between 30B and 50B in a funding round that would value it at up to $950B (Mike Isaac/New York Times)

Model ReleasesDGX agent

Mike Isaac / New York Times: Sources: Anthropic is in talks to raise between 30B and 50B in a funding round that would value it at up to 950B — The start-up, which recently released a powerful A.I. mo

SpaceX is officially my fav art account now

Model ReleasesDGX agent

SpaceX's official social media account is being praised for its engaging visual content and aesthetic presentation, earning recognition as a favorite among followers for its artistic approach to shari

Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers

Model ReleasesDGX agent

arXiv:2605.09403v1 Announce Type: cross Abstract: Architectural choices inside the Transformer feedforward network (FFN) block do not merely affect the block itself; they reshape the computations lear

SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?

Model ReleasesDGX agent

arXiv:2602.03916v3 Announce Type: replace-cross Abstract: Spatial reasoning is a fundamental aspect of human cognition, yet it remains a major challenge for contemporary vision-language models (VLMs).

SPDEBench: An Extensive Benchmark for Learning Stochastic PDEs

Model ReleasesDGX agent

arXiv:2505.18511v2 Announce Type: replace Abstract: Stochastic Partial Differential Equations (SPDEs) driven by random noise play a central role in modeling physical processes with rough spatio-tempor

Spectral Characterization and Mitigation of Sequential Knowledge Editing Collapse

Model ReleasesDGX agent

arXiv:2601.11042v2 Announce Type: replace-cross Abstract: Sequential knowledge editing in large language models often causes catastrophic collapse of the model's general abilities, especially for para

SpectraLLM: Uncovering the Ability of LLMs for Molecular Structure Elucidation from Multi-Spectral Data

Model ReleasesDGX agent

arXiv:2508.08441v3 Announce Type: replace-cross Abstract: Automated molecular structure elucidation remains challenging, as existing approaches often depend on pre-compiled databases or restrict thems

Spherical Boltzmann machines: a solvable theory of learning and generation in energy-based models

Model ReleasesDGX agent

arXiv:2605.09031v1 Announce Type: new Abstract: Energy-based models (EBMs) are flexible generative architectures inspired by statistical physics, but their learning and generative properties remain po

Statistical Model Checking of the Keynes+Schumpeter Model: A Transient Sensitivity Analysis of a Macroeconomic ABM

Model ReleasesDGX agent

arXiv:2605.10447v1 Announce Type: cross Abstract: Agent-based models (ABMs) are increasingly used in macroeconomics, but their analysis still often relies on ad hoc Monte Carlo campaigns with heteroge

Statistical Scouting Finds Debate-Safe but Not Debate-Useful Cases: A Matched-Ceiling Study of Open-Weight LLM Reasoning Protocols

Model ReleasesDGX agent

arXiv:2605.09618v1 Announce Type: new Abstract: When should a language model answer directly, sample and vote, or engage in multi-agent debate? Recent work shows voting often explains much of the gain

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens

Model ReleasesDGX agent

arXiv:2604.02608v2 Announce Type: replace Abstract: Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both stee

Step Rejection Fine-Tuning: A Practical Distillation Recipe

Model ReleasesDGX agent

arXiv:2605.10674v1 Announce Type: cross Abstract: Rejection Fine-Tuning (RFT) is a standard method for training LLM agents, where unsuccessful trajectories are discarded from the training set. In the

Strategic commitments shape collective cybersecurity under AI inequality

Model ReleasesDGX agent

arXiv:2605.09415v1 Announce Type: new Abstract: The growing integration of AI into cybersecurity is reshaping the balance between attackers and defenders. When access to advanced AI-enabled defence to

Strategic Exploitation in LLM Agent Markets: A Simulation Framework for E-Commerce Trust

Model ReleasesDGX agent

arXiv:2605.10059v1 Announce Type: new Abstract: Agent-based modeling (ABM) has long been used in economics to study human behavior, and large language model (LLM) agents now enable new forms of social

Structure-Preserving Reconstruction of Convex Lipschitz Functionals on Hilbert Spaces from Finite Samples

Model ReleasesDGX agent

arXiv:2605.08559v1 Announce Type: cross Abstract: Convex functionals are ubiquitous in applied analysis, appearing as value functions, risk measures, super-hedging prices, and loss functionals in mach

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

Model ReleasesDGX agent

arXiv:2605.08366v1 Announce Type: new Abstract: We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test W

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

Model ReleasesDGX agent

arXiv:2605.08412v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

Model ReleasesDGX agent

arXiv:2605.10129v1 Announce Type: new Abstract: Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and u

TabPFN-3 just released: a pre-trained tabular foundation model for up to 1M rows [R][N]

Model ReleasesDGX agent

TabPFN-3 is a pre-trained tabular foundation model that supports datasets up to 1,000,000 rows × 200 features , representing a significant scaling improvement for the TabPFN family. The model delivers

Tabular Foundation Model for Generative Modelling

Model ReleasesDGX agent

arXiv:2605.09424v1 Announce Type: new Abstract: Generative modelling is a demanding test of foundation models, because it requires robust, holistic representation learning for a given data modality, r

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.09539v1 Announce Type: new Abstract: Multi-agent systems (MAS) have emerged as a promising paradigm for solving complex tasks. Recent work has explored self-evolving MAS that automatically

Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation

Model ReleasesDGX agent

arXiv:2505.11604v5 Announce Type: replace Abstract: Editing presentation slides is a frequent yet tedious task, ranging from creative layout design to repetitive text maintenance. While recent GUI-bas

Teaching LLMs to See Graphs: Unifying Text and Structural Reasoning

Model ReleasesDGX agent

arXiv:2605.10247v1 Announce Type: new Abstract: Using Large Language Models (LLMs) to process graph-structured data is an active research area, yet current state-of-the-art approaches typically rely o

TeleResilienceBench: Quantifying Resilience for LLM Reasoning in Telecommunications

Model ReleasesDGX agent

arXiv:2605.09929v1 Announce Type: new Abstract: Deploying large language models in telecommunications requires more than task accuracy. In realistic workflows, a model may inherit partially completed

'tell me about recent funding and news for openai and anthropic' result with @wokeloai mcp included today's trial, tomoro acquisition, and c…

Model ReleasesDGX agent

Yohei Nakajima shared recent funding and news updates about OpenAI and Anthropic on X, including information about a trial involving the wokeloai MCP, a tomorrow acquisition, and additional details (c

Test-Time Speculation

Model ReleasesDGX agent

arXiv:2605.09329v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by using a fast draft model to generate tokens and a more accurate target model to verify them. Its perfo

Text-Guided Multi-Scale Frequency Representation Adaptation

Model ReleasesDGX agent

arXiv:2605.08181v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning methods introduce a small number of training parameters, enabling pre-trained models to adapt rapidly to new data dist

TFM-Retouche: A Lightweight Input-Space Adapter for Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2605.06047v2 Announce Type: replace-cross Abstract: Tabular foundation models (TFMs), such as TabPFN-2.6, TabICLv2, ConTextTab, Mitra, LimiX, and TabDPT, achieve strong zero-shot performance thr

The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection

Model ReleasesDGX agent

arXiv:2605.10334v1 Announce Type: new Abstract: Recent deepfake detection methods demonstrate improved cross-dataset generalization, yet the underlying mechanisms remain underexplored. We introduce th

The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

Model ReleasesDGX agent

arXiv:2605.08427v1 Announce Type: new Abstract: Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in

The autoPET3 Challenge: Automated Lesion Segmentation in Whole-Body PET/CT nicode{x2013} Multitracer Multicenter Generalization

Model ReleasesDGX agent

arXiv:2605.05775v2 Announce Type: replace-cross Abstract: We report the design and results of the third autoPET challenge (MICCAI 2024), which benchmarked automated lesion segmentation in whole-body P

The Differences Between Direct Alignment Algorithms are a Blur

Model ReleasesDGX agent

arXiv:2502.01237v3 Announce Type: replace Abstract: Direct Alignment Algorithms (DAAs) simplify LLM alignment by directly optimizing policies, bypassing reward modeling and RL. While DAAs differ in th

The Echo Amplifies the Knowledge: Somatic Marker Analogues in Language Models via Emotion Vector Re-Injection

Model ReleasesDGX agent

arXiv:2605.08611v1 Announce Type: new Abstract: Current language model memory systems store what happened but not how it felt. This distinction -- between semantic memory (knowing about a past event)

The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs

Model ReleasesDGX agent

arXiv:2605.08737v1 Announce Type: cross Abstract: On-policy distillation (OPD) is widely used for LLM post-training. When pushed with a reward-extrapolation coefficient lambda > 1, the student can lif

The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws

Model ReleasesDGX agent

arXiv:2605.09887v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) operationalise the linear representation hypothesis: they reconstruct model activations as sparse linear combinations of in

The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations

Model ReleasesDGX agent

arXiv:2605.09195v1 Announce Type: new Abstract: Large language models confidently produce outdated answers, and no existing method can detect them. We show this is not an engineering failure but a str

The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning

Model ReleasesDGX agent

arXiv:2605.08746v1 Announce Type: new Abstract: In training a neural network with gradient descent (GD), each iteration induces a linear operator that governs first-order updates to a model's internal

The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark

Model ReleasesDGX agent

arXiv:2605.09900v1 Announce Type: new Abstract: A vision-language model can look at a knot diagram and report what it sees, yet fail to act on that structure. KnotBench pairs an 858,318-image corpus f

The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies

Model ReleasesDGX agent

arXiv:2605.10799v1 Announce Type: cross Abstract: Corruption studies, the primary tool for evaluating chain-of-thought (CoT) faithfulness, identify which chain positions are 'computationally important

The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs

Model ReleasesDGX agent

arXiv:2605.09844v1 Announce Type: new Abstract: The Metacognitive Probe is an exploratory five-task, 15-slot diagnostic that decomposes an LLM's confidence behaviour into five behaviourally-distinct d

The Realignment Problem: When Right becomes Wrong in LLMs

Model ReleasesDGX agent

arXiv:2511.02623v2 Announce Type: replace Abstract: Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over tim

The silent removal of Study Mode from ChatGPT is a big mistake (both Claude and Gemini still have theirs) We have enough evidence that using…

Model ReleasesDGX agent

The silent removal of Study Mode from ChatGPT is a big mistake (both Claude and Gemini still have theirs) We have enough evidence that using AI in assistant mode to study can hurt learning because it

The Silent Vote: Improving Zero-Shot LLM Reliability by Aggregating Semantic Neighborhoods

Model ReleasesDGX agent

arXiv:2605.09739v1 Announce Type: cross Abstract: Large Language Models are increasingly used as zero-shot classifiers in complex reasoning tasks. However, standard constrained decoding suffers from a

The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory

Model ReleasesDGX agent

arXiv:2605.09330v1 Announce Type: cross Abstract: Agentic memory enables LLMs to persist information beyond a single context window and reuse it in later decisions, but it also introduces a new vulner

The US House Oversight Committee launches a probe into potential conflicts in Sam Altman's personal investments; letter: several GOP AGs call for an SEC review (Wall Street Journal)

Model ReleasesDGX agent

Wall Street Journal: The US House Oversight Committee launches a probe into potential conflicts in Sam Altman's personal investments; letter: several GOP AGs call for an SEC review — Republican-led Ho

The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models

Model ReleasesDGX agent

arXiv:2601.02954v3 Announce Type: replace-cross Abstract: Large audio-language models have made rapid progress in recognizing what is present in an audio clip, but spatial audio-language understanding

The Wristband Gaussian Loss: Deterministic, Composable Latents via a Sphere-Interval Decomposition

Model ReleasesDGX agent

arXiv:2605.08749v1 Announce Type: new Abstract: We present the Wristband Gaussian Loss, a deterministic batch loss for Gaussianizing point embeddings without sampling, KL terms, or iterative transport

This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization,…

Model ReleasesDGX agent

This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization, custom kernels, and rack-scale NVLink turn GB200 into faste

ThreatCore: A Benchmark for Explicit and Implicit Threat Detection

Model ReleasesDGX agent

arXiv:2605.10563v1 Announce Type: cross Abstract: Threat detection in Natural Language Processing lacks consistent definitions and standardized benchmarks, and is often conflated with broader phenomen

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning

Model ReleasesDGX agent

arXiv:2605.09544v1 Announce Type: new Abstract: Tool-integrated reasoning has emerged as a promising paradigm for enhancing large language models with external computation, retrieval, and execution ca

← Previous
1…268269270271272…377
Next →