AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,585 results
Model Releases

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens

DGX agent

arXiv:2604.02608v2 Announce Type: replace Abstract: Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both stee

model-releasesarxiv-cs-lg
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Step Rejection Fine-Tuning: A Practical Distillation Recipe

DGX agent

arXiv:2605.10674v1 Announce Type: cross Abstract: Rejection Fine-Tuning (RFT) is a standard method for training LLM agents, where unsuccessful trajectories are discarded from the training set. In the

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Strategic commitments shape collective cybersecurity under AI inequality

DGX agent

arXiv:2605.09415v1 Announce Type: new Abstract: The growing integration of AI into cybersecurity is reshaping the balance between attackers and defenders. When access to advanced AI-enabled defence to

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Strategic Exploitation in LLM Agent Markets: A Simulation Framework for E-Commerce Trust

DGX agent

arXiv:2605.10059v1 Announce Type: new Abstract: Agent-based modeling (ABM) has long been used in economics to study human behavior, and large language model (LLM) agents now enable new forms of social

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Structure-Preserving Reconstruction of Convex Lipschitz Functionals on Hilbert Spaces from Finite Samples

DGX agent

arXiv:2605.08559v1 Announce Type: cross Abstract: Convex functionals are ubiquitous in applied analysis, appearing as value functions, risk measures, super-hedging prices, and loss functionals in mach

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

DGX agent

arXiv:2605.08366v1 Announce Type: new Abstract: We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test W

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

DGX agent

arXiv:2605.08412v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

DGX agent

arXiv:2605.10129v1 Announce Type: new Abstract: Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and u

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

TabPFN-3 just released: a pre-trained tabular foundation model for up to 1M rows [R][N]

DGX agent

TabPFN-3 is a pre-trained tabular foundation model that supports datasets up to 1,000,000 rows × 200 features , representing a significant scaling improvement for the TabPFN family. The model delivers

model-releasesr-machinelearning
12 May 2026
Model Releases

Tabular Foundation Model for Generative Modelling

DGX agent

arXiv:2605.09424v1 Announce Type: new Abstract: Generative modelling is a demanding test of foundation models, because it requires robust, holistic representation learning for a given data modality, r

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems

DGX agent

arXiv:2605.09539v1 Announce Type: new Abstract: Multi-agent systems (MAS) have emerged as a promising paradigm for solving complex tasks. Recent work has explored self-evolving MAS that automatically

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation

DGX agent

arXiv:2505.11604v5 Announce Type: replace Abstract: Editing presentation slides is a frequent yet tedious task, ranging from creative layout design to repetitive text maintenance. While recent GUI-bas

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Teaching LLMs to See Graphs: Unifying Text and Structural Reasoning

DGX agent

arXiv:2605.10247v1 Announce Type: new Abstract: Using Large Language Models (LLMs) to process graph-structured data is an active research area, yet current state-of-the-art approaches typically rely o

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

TeleResilienceBench: Quantifying Resilience for LLM Reasoning in Telecommunications

DGX agent

arXiv:2605.09929v1 Announce Type: new Abstract: Deploying large language models in telecommunications requires more than task accuracy. In realistic workflows, a model may inherit partially completed

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

'tell me about recent funding and news for openai and anthropic' result with @wokeloai mcp included today's trial, tomoro acquisition, and c…

DGX agent

Yohei Nakajima shared recent funding and news updates about OpenAI and Anthropic on X, including information about a trial involving the wokeloai MCP, a tomorrow acquisition, and additional details (c

model-releasesyohei-nakajima--x
12 May 2026
Model Releases

Test-Time Speculation

DGX agent

arXiv:2605.09329v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by using a fast draft model to generate tokens and a more accurate target model to verify them. Its perfo

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Text-Guided Multi-Scale Frequency Representation Adaptation

DGX agent

arXiv:2605.08181v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning methods introduce a small number of training parameters, enabling pre-trained models to adapt rapidly to new data dist

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TFM-Retouche: A Lightweight Input-Space Adapter for Tabular Foundation Models

DGX agent

arXiv:2605.06047v2 Announce Type: replace-cross Abstract: Tabular foundation models (TFMs), such as TabPFN-2.6, TabICLv2, ConTextTab, Mitra, LimiX, and TabDPT, achieve strong zero-shot performance thr

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection

DGX agent

arXiv:2605.10334v1 Announce Type: new Abstract: Recent deepfake detection methods demonstrate improved cross-dataset generalization, yet the underlying mechanisms remain underexplored. We introduce th

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

DGX agent

arXiv:2605.08427v1 Announce Type: new Abstract: Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The autoPET3 Challenge: Automated Lesion Segmentation in Whole-Body PET/CT nicode{x2013} Multitracer Multicenter Generalization

DGX agent

arXiv:2605.05775v2 Announce Type: replace-cross Abstract: We report the design and results of the third autoPET challenge (MICCAI 2024), which benchmarked automated lesion segmentation in whole-body P

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Differences Between Direct Alignment Algorithms are a Blur

DGX agent

arXiv:2502.01237v3 Announce Type: replace Abstract: Direct Alignment Algorithms (DAAs) simplify LLM alignment by directly optimizing policies, bypassing reward modeling and RL. While DAAs differ in th

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

The Echo Amplifies the Knowledge: Somatic Marker Analogues in Language Models via Emotion Vector Re-Injection

DGX agent

arXiv:2605.08611v1 Announce Type: new Abstract: Current language model memory systems store what happened but not how it felt. This distinction -- between semantic memory (knowing about a past event)

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs

DGX agent

arXiv:2605.08737v1 Announce Type: cross Abstract: On-policy distillation (OPD) is widely used for LLM post-training. When pushed with a reward-extrapolation coefficient lambda > 1, the student can lif

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws

DGX agent

arXiv:2605.09887v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) operationalise the linear representation hypothesis: they reconstruct model activations as sparse linear combinations of in

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations

DGX agent

arXiv:2605.09195v1 Announce Type: new Abstract: Large language models confidently produce outdated answers, and no existing method can detect them. We show this is not an engineering failure but a str

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning

DGX agent

arXiv:2605.08746v1 Announce Type: new Abstract: In training a neural network with gradient descent (GD), each iteration induces a linear operator that governs first-order updates to a model's internal

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark

DGX agent

arXiv:2605.09900v1 Announce Type: new Abstract: A vision-language model can look at a knot diagram and report what it sees, yet fail to act on that structure. KnotBench pairs an 858,318-image corpus f

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies

DGX agent

arXiv:2605.10799v1 Announce Type: cross Abstract: Corruption studies, the primary tool for evaluating chain-of-thought (CoT) faithfulness, identify which chain positions are 'computationally important

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs

DGX agent

arXiv:2605.09844v1 Announce Type: new Abstract: The Metacognitive Probe is an exploratory five-task, 15-slot diagnostic that decomposes an LLM's confidence behaviour into five behaviourally-distinct d

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Realignment Problem: When Right becomes Wrong in LLMs

DGX agent

arXiv:2511.02623v2 Announce Type: replace Abstract: Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over tim

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

The silent removal of Study Mode from ChatGPT is a big mistake (both Claude and Gemini still have theirs) We have enough evidence that using…

DGX agent

The silent removal of Study Mode from ChatGPT is a big mistake (both Claude and Gemini still have theirs) We have enough evidence that using AI in assistant mode to study can hurt learning because it

model-releasesethan-mollick--x
12 May 2026
Model Releases

The Silent Vote: Improving Zero-Shot LLM Reliability by Aggregating Semantic Neighborhoods

DGX agent

arXiv:2605.09739v1 Announce Type: cross Abstract: Large Language Models are increasingly used as zero-shot classifiers in complex reasoning tasks. However, standard constrained decoding suffers from a

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory

DGX agent

arXiv:2605.09330v1 Announce Type: cross Abstract: Agentic memory enables LLMs to persist information beyond a single context window and reuse it in later decisions, but it also introduces a new vulner

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The US House Oversight Committee launches a probe into potential conflicts in Sam Altman's personal investments; letter: several GOP AGs call for an SEC review (Wall Street Journal)

DGX agent

Wall Street Journal: The US House Oversight Committee launches a probe into potential conflicts in Sam Altman's personal investments; letter: several GOP AGs call for an SEC review — Republican-led Ho

model-releasestechmeme
12 May 2026
Model Releases

The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models

DGX agent

arXiv:2601.02954v3 Announce Type: replace-cross Abstract: Large audio-language models have made rapid progress in recognizing what is present in an audio clip, but spatial audio-language understanding

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Wristband Gaussian Loss: Deterministic, Composable Latents via a Sphere-Interval Decomposition

DGX agent

arXiv:2605.08749v1 Announce Type: new Abstract: We present the Wristband Gaussian Loss, a deterministic batch loss for Gaussianizing point embeddings without sampling, KL terms, or iterative transport

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization,…

DGX agent

This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization, custom kernels, and rack-scale NVLink turn GB200 into faste

model-releasesperplexity--x
12 May 2026
Model Releases

ThreatCore: A Benchmark for Explicit and Implicit Threat Detection

DGX agent

arXiv:2605.10563v1 Announce Type: cross Abstract: Threat detection in Natural Language Processing lacks consistent definitions and standardized benchmarks, and is often conflated with broader phenomen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning

DGX agent

arXiv:2605.09544v1 Announce Type: new Abstract: Tool-integrated reasoning has emerged as a promising paradigm for enhancing large language models with external computation, retrieval, and execution ca

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TIDES: Implicit Time-Awareness in Selective State Space Models

DGX agent

arXiv:2605.09742v1 Announce Type: cross Abstract: Selective state space models (SSMs), such as Mamba, achieve strong per-token expressivity by making the time discretization step Tilde{Delta} a learne

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Tight Generalization Bounds for Noiseless Inverse Optimization

DGX agent

arXiv:2605.08866v1 Announce Type: cross Abstract: Inverse optimization (IO) seeks to infer the parameters of a decision-maker's objective from observed context--action data. We study noiseless IO, whe

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

TiledAttention: a CUDA Tile SDPA Kernel for PyTorch

DGX agent

arXiv:2603.01960v2 Announce Type: replace-cross Abstract: TiledAttention is a scaled dot-product attention (SDPA) forward operator for SDPA research on NVIDIA GPUs. Implemented in cuTile Python (TileI

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TINS: Test-time ID-prototype-separated Negative Semantics Learning for OOD Detection

DGX agent

arXiv:2605.10756v1 Announce Type: new Abstract: Vision-language models enable OOD detection by comparing image alignment with ID labels and negative semantics. Existing negative-label-based methods ma

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

To Redact, or not to Redact? A Local LLM Approach to Deliberative Process Privilege Classification

DGX agent

arXiv:2605.10211v1 Announce Type: cross Abstract: Government transparency laws, like the Freedom of Information (FOIA) acts in the United States and United Kingdom, and the Woo (Open Government Act) i

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models

DGX agent

arXiv:2605.09904v1 Announce Type: new Abstract: Video large language models (Video-LLMs) have achieved remarkable progress in general video understanding, yet their ability to maintain temporal object

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation

DGX agent

arXiv:2605.08541v1 Announce Type: new Abstract: Neural scaling laws approximate a language model's loss as a power-law function of parameter count N and token count D. Following Chinchilla-style compu

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Towards a Large Language-Vision Question Answering Model for MSTAR Automatic Target Recognition

DGX agent

arXiv:2605.10772v1 Announce Type: cross Abstract: Large language-vision models (LLVM), such as OpenAI's ChatGPT and GPT-4, have gained prominence as powerful tools for analyzing text and imagery. The

model-releasesarxiv-cs-ai
12 May 2026
← Previous
1…336337338339340…471
Next →