AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
12 May 2026

Re^2Math: Benchmarking Theorem Retrieval in Research-Level Mathematics

Model ReleasesDGX agent

arXiv:2605.09012v1 Announce Type: new Abstract: Large language models are increasingly capable at closed-world mathematical reasoning, but research assistance also requires source-grounded use of the

Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation

Model ReleasesDGX agent

arXiv:2605.10430v1 Announce Type: cross Abstract: Estimating heterogeneous treatment effects with machine learning has attracted substantial attention in both academic research and industrial practice

Really amazing dissection. LLMs hitting rarely. Tools list is amazing.

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Really amazing dissection. LLMs hitting rarely. Tools list is amazing. 🤩🤯🤩 Claude Code (still not AGI but biggest advance since GPT-4) is the most neurosymbolic thing I have ever seen in my life. 53 s

ReaMOT: A Benchmark and Framework for Reasoning-based Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2505.20381v4 Announce Type: replace Abstract: Referring Multi-Object Tracking (RMOT) aims to track targets specified by language instructions. However, existing RMOT paradigms heavily rely on ex

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

Model ReleasesDGX agent

arXiv:2604.01527v3 Announce Type: replace-cross Abstract: Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fi

Reasoning emerges from constrained inference manifolds in large language models

Model ReleasesDGX agent

arXiv:2605.08142v1 Announce Type: cross Abstract: Reasoning in large language models is predominantly evaluated through labeled benchmarks, conflating task performance with the quality of internal inf

Recovering Physical Dynamics from Discrete Observations via Intrinsic Differential Consistency

Model ReleasesDGX agent

arXiv:2605.08454v1 Announce Type: cross Abstract: Recovering continuous-time dynamics from discrete observations is difficult because local supervision (e.g., pointwise regression targets, derivative

Recursive Language Models

Model ReleasesDGX agent

arXiv:2512.24601v3 Announce Type: replace Abstract: We study allowing large language models (LLMs) to process arbitrarily long prompts through the lens of inference-time scaling. We propose Recursive

REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?

Model ReleasesDGX agent

arXiv:2505.10872v4 Announce Type: replace-cross Abstract: Robot task planning decomposes human instructions into executable action sequences that enable robots to complete a series of complex tasks. A

Reinforcement Learning Measurement Model

Model ReleasesDGX agent

arXiv:2605.09305v1 Announce Type: cross Abstract: Interactive assessments generate sequential process data that are not well handled by conventional item response models. Existing MDP-based measuremen

Relative Kinetic Utility for Reasoning-Aware Structural Pruning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.09008v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) prompting symbolized a huge improvement of reasoning capabilities of Large Language Models (LLMs). However, scaling up test-tim

RelBench v2: A Large-Scale Benchmark and Repository for Relational Data

Model ReleasesDGX agent

arXiv:2602.12606v2 Announce Type: replace Abstract: Relational deep learning (RDL) has emerged as a powerful paradigm for learning directly on relational databases by modeling entities and their relat

Reliable LLM-Based Edge-Cloud-Expert Cascades for Telecom Knowledge Systems

Model ReleasesDGX agent

arXiv:2512.20012v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are emerging as key enablers of automation in domains such as telecommunications, assisting with tasks including

Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction

Model ReleasesDGX agent

arXiv:2605.08871v1 Announce Type: cross Abstract: Large-scale machine learning models are trained on clusters of machines that exhibit heterogeneous performance due to hardware variability, network de

ReorgGS: Equivalent Distribution Reorganization for 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2605.08739v1 Announce Type: new Abstract: A converged 3D Gaussian Splatting (3DGS) model may approximate the target scene while remaining poorly parameterized for further optimization. We identi

Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs

Model ReleasesDGX agent

arXiv:2605.09239v1 Announce Type: new Abstract: Large language models fail at counting repeated tokens despite strong performance on broader reasoning benchmarks. These failures are commonly attribute

ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions

Model ReleasesDGX agent

arXiv:2605.08197v1 Announce Type: cross Abstract: Most causal benchmarks for language models score local answers or graph structure. We introduce ReplaySCM, a 1,300 item benchmark for executable causa

Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

Model ReleasesDGX agent

arXiv:2605.10848v1 Announce Type: cross Abstract: Does a lexical retriever suffice as large language models (LLMs) become more capable in an agentic loop? This question naturally arises when building

Rethinking Random Transformers as Adaptive Sequence Smoothers for Sleep Staging

Model ReleasesDGX agent

arXiv:2605.09905v1 Announce Type: cross Abstract: Automatic sleep staging commonly adopts Transformers under the assumption that they learn complex long-range dependencies. We challenge this view by r

Retrieval Mechanisms Surpass Long-Context Scaling in Time Series Forecasting

Model ReleasesDGX agent

arXiv:2605.08217v1 Announce Type: new Abstract: Time Series Foundation Models (TSFMs) have borrowed the long context paradigm from natural language processing under the premise that feeding more histo

Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs

Model ReleasesDGX agent

arXiv:2605.10094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models show strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades u

RewardHarness: Self-Evolving Agentic Post-Training

Model ReleasesDGX agent

arXiv:2605.08703v1 Announce Type: new Abstract: Evaluating instruction-guided image edits requires rewards that reflect subtle human preferences, yet current reward models typically depend on large-sc

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

Model ReleasesDGX agent

arXiv:2605.10921v1 Announce Type: new Abstract: Memory is a critical component of robotic intelligence, as robots must rely on past observations and actions to accomplish long-horizon tasks in partial

Robust Server Defense Against Unreliable Clients in One-Shot Fair Collaborative Machine Learning

Model ReleasesDGX agent

arXiv:2605.08616v1 Announce Type: new Abstract: Collaborative machine learning (CML) enables multiple clients to train a global model jointly in a data-distributed setting. To address data privacy and

Robust Spectral Watermark for Synthetic Tabular Data

Model ReleasesDGX agent

arXiv:2511.21600v2 Announce Type: replace-cross Abstract: The rise of generative AI has enabled the production of high-fidelity synthetic tabular data across fields such as healthcare, finance, and pu

ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention

Model ReleasesDGX agent

arXiv:2603.22016v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant verifi

Route Before Retrieve: Activating Latent Routing Abilities of LLMs for RAG vs. Long-Context Selection

Model ReleasesDGX agent

arXiv:2605.10235v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have expanded the context window to beyond 128K tokens, enabling long-document understanding and multi-s

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement

Model ReleasesDGX agent

arXiv:2605.09730v1 Announce Type: new Abstract: Iterative self-refinement is a popular inference-time reliability technique, but its effectiveness in code-mode tool use depends heavily on the structur

RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild

Model ReleasesDGX agent

arXiv:2605.10357v1 Announce Type: cross Abstract: Multimodal misinformation increasingly leverages visual persuasion, where repurposed or manipulated images strengthen misleading text. We introduce ex

S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum Domain

Model ReleasesDGX agent

arXiv:2605.08589v1 Announce Type: new Abstract: Parameter Efficient Fine-Tuning (PEFT) is a key technique for adapting a large pretrained model to downstream tasks by fine-tuning only a small number o

SACHI: Structured Agent Coordination via Holistic Information Integration in Multi-Agent Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.08391v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning agents that act on partial local observations face a fundamental information bottleneck: the knowledge ne

SAFA-SNN: Sparsity-Aware On-Device Few-Shot Class-Incremental Learning with Fast-Adaptive Structure of Spiking Neural Network

Model ReleasesDGX agent

arXiv:2510.03648v2 Announce Type: replace Abstract: Continuous learning of novel classes is crucial for edge devices to preserve data privacy and maintain reliable performance in dynamic environments.

SAP SAPPHIRE 2026: Google Cloud unveils unified agentic vision and massive compute scaling

Model ReleasesDGX agent

In today's hyper-connected market, an enterprise's most valuable asset — mission-critical data — often remains trapped in legacy silos. For years, leadership teams have navigated a data pipeline dilem

SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?

Model ReleasesDGX agent

arXiv:2602.00327v2 Announce Type: replace Abstract: We explore the use of large language models (LLMs) for next-utterance anticipation in human dialogue. Despite recent advances in LLMs demonstrating

Scalable Gaussian process inference via neural feature maps

Model ReleasesDGX agent

arXiv:2605.10285v1 Announce Type: cross Abstract: We present a theoretically grounded Gaussian process framework that leverages neural feature maps to construct expressive kernels. We show that the le

SCALAR: A Neurosymbolic Framework for Automated Conjecture and Reasoning in Quantum Circuit Analysis

Model ReleasesDGX agent

arXiv:2605.10327v1 Announce Type: cross Abstract: In this paper, we present SCALAR (Symbolic Conjecture and LLM-Assisted Reasoning), a neurosymbolic framework for automated conjecture generation in qu

Scaling Limits of Long-Context Transformers

Model ReleasesDGX agent

arXiv:2605.08505v1 Announce Type: cross Abstract: We study the long-context limit of softmax self-attention with a fixed query and a random context of n i.i.d. keys on the sphere, viewing the inverse

Scaling the Memory of Balanced Adam

Model ReleasesDGX agent

arXiv:2605.10119v1 Announce Type: new Abstract: Recent evidence suggests that Adam performs robustly when its momentum parameters are tied, eta_1=eta_2, reducing the optimizer to a single remaining pa

Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality

Model ReleasesDGX agent

arXiv:2605.10142v1 Announce Type: cross Abstract: Artificial intelligence models are increasingly scaled to improve predictive accuracy, yet it remains unclear whether scale improves the quality of po

Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs

Model ReleasesDGX agent

arXiv:2509.02372v3 Announce Type: replace-cross Abstract: Large Language Models have become critical to modern software development, but their reliance on uncurated web-scale datasets for training int

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems

Model ReleasesDGX agent

arXiv:2605.10246v1 Announce Type: new Abstract: AI scientist systems are increasingly deployed for autonomous research, yet their academic integrity has never been systematically evaluated. We introdu

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation

Model ReleasesDGX agent

arXiv:2605.10187v1 Announce Type: new Abstract: Scientific reasoning is a key aspect of human intelligence, requiring the integration of multimodal inputs, domain expertise, and multi-step inference a

SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness

Model ReleasesDGX agent

arXiv:2603.14889v2 Announce Type: replace-cross Abstract: The rapid evolution of end-to-end spoken dialogue systems demands transcending mere textual semantics to incorporate paralinguistic nuances an

SeBA: Semi-supervised few-shot learning via Separated-at-Birth Alignment for tabular data

Model ReleasesDGX agent

arXiv:2605.08519v1 Announce Type: new Abstract: Learning from scarce labeled data with a larger pool of unlabeled samples, known as semi-supervised few-shot learning (SS-FSL), remains critical for app

Seed Hijacking of LLM Sampling and Quantum Random Number Defense

Model ReleasesDGX agent

arXiv:2605.08313v1 Announce Type: cross Abstract: Large language models (LLMs) rely on deterministic pseudorandom number generators (PRNGs) for autoregressive sampling, creating a critical supply-chai

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning

Model ReleasesDGX agent

arXiv:2605.09266v1 Announce Type: new Abstract: We introduce SeePhys Pro, a fine-grained modality transfer benchmark that studies whether models preserve the same reasoning capability when critical in

Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind

Model ReleasesDGX agent

arXiv:2603.26089v2 Announce Type: replace-cross Abstract: The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind

Selective LoRA for Visual Tokens and Attention Heads

Model ReleasesDGX agent

arXiv:2512.19219v2 Announce Type: replace-cross Abstract: Low-rank adaptation (LoRA) is widely used for parameter-efficient fine-tuning, but its standard all-token, all-head design ignores the heterog

SEMASIA: A Large-Scale Dataset of Semantically Structured Latent Representations

Model ReleasesDGX agent

arXiv:2605.09485v1 Announce Type: new Abstract: Latent representations learned by neural networks often exhibit semantic structure, where concept similarity is reflected by geometric proximity in embe

Semi-Supervised Neural Super-Resolution for Mesh-Based Simulations

Model ReleasesDGX agent

arXiv:2605.09284v1 Announce Type: cross Abstract: Mesh-based simulations provide high-fidelity solutions to partial differential equations (PDEs), but achieving such accuracy typically requires fine m

Sens-VisualNews: A Benchmark Dataset for Sensational Image Detection

Model ReleasesDGX agent

arXiv:2605.10394v1 Announce Type: new Abstract: The detection of sensational content in media items can be a critical filtering mechanism for identifying check-worthy content and flagging potential di

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.10576v1 Announce Type: cross Abstract: Low-level visual perception underpins reliable remote sensing (RS) image analysis, yet current image quality assessment (IQA) methods output uninterpr

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

Model ReleasesDGX agent

arXiv:2605.09906v1 Announce Type: new Abstract: Audio and vision provide complementary evidence for audio-visual question answering, yet current audio-visual large language models may suffer from cros

Sequential Causal Discovery with Noisy Language Model Priors

Model ReleasesDGX agent

arXiv:2506.16234v2 Announce Type: replace Abstract: Causal discovery from observational data typically assumes access to complete data and availability of perfect domain experts. In practice, data oft

Sequential Feature Selection for Efficient Landslide Segmentation from Multi-Spectral Data

Model ReleasesDGX agent

arXiv:2605.09746v1 Announce Type: cross Abstract: Landslide detection from satellite imagery has advanced through deep learning, yet most models rely on large, highly correlated spectral-topographic i

Set Prediction for Next-Day Active Fire Forecasting

Model ReleasesDGX agent

arXiv:2605.10298v1 Announce Type: new Abstract: Accurate next-day active fire forecasts can support early warning, disaster response, forest risk assessment, and downstream estimation of fire-related

simpleposter: a simple baseline for product poster generation

Model ReleasesDGX agent

arXiv:2605.08784v1 Announce Type: new Abstract: Product poster generation poses distinct challenges beyond general poster design, requiring both faithful preservation of product appearance and precise

Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data

Model ReleasesDGX agent

arXiv:2605.10498v1 Announce Type: cross Abstract: Long-tailed distributions in class-imbalanced data present a fundamental challenge for deep learning models, which tend to be biased toward majority c

Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success

Model ReleasesDGX agent

arXiv:2605.09070v1 Announce Type: cross Abstract: Many jailbreak attack research papers report attack success rates for a limited number of parameter settings, even though there are many combinations

Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data Loaders

Model ReleasesDGX agent

arXiv:2605.08731v1 Announce Type: cross Abstract: JPEG decode is routine ML infrastructure, but Python decoder choices are often justified by single-process, single-thread microbenchmarks. We audit th

← Previous
1…267268269270271…377
Next →