AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

86,965Total entries
1Added by human
86,964Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,458 results
12 May 2026

Generating Leakage-Free Benchmarks for Robust RAG Evaluation

Model ReleasesDGX agent

arXiv:2605.08838v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is widely used to augment large language models (LLMs) with external knowledge. However, many benchmark datasets,

Geometry-free prediction of inertial lift forces in microfluidic devices using deep learning

Model ReleasesDGX agent

arXiv:2605.08109v1 Announce Type: new Abstract: Inertial microfluidic devices (IMDs) offer low-cost, high-throughput alternative techniques for many traditional particle- (or cell-) manipulation tasks

Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.10893v1 Announce Type: new Abstract: Large vision-language models suffer from visual ungroundedness: they can produce a fluent, confident, and even correct response driven entirely by langu

HS-FNO: History-Space Fourier Neural Operator for Non-Markovian Partial Differential Equations

Model ReleasesDGX agent

arXiv:2605.09523v1 Announce Type: new Abstract: Neural operators provide fast surrogate models for time-dependent partial differential equations, but their standard autoregressive use usually assumes

In-Context Fixation: When Demonstrated Labels Override Semantics in Few-Shot Classification

Model ReleasesDGX agent

arXiv:2605.08295v1 Announce Type: cross Abstract: While random demonstration labels barely hurt in-context learning (Min et al., 2022), we show that homogeneous labels--even semantically valid ones--c

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

ApplicationsDGX agent

arXiv:2601.22143v2 Announce Type: replace-cross Abstract: Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented abilit

LEVI: Stronger Search Architectures Can Substitute for Larger LLMs in Evolutionary Search

Model ReleasesDGX agent

arXiv:2605.09764v1 Announce Type: cross Abstract: LLM-guided evolutionary methods such as AlphaEvolve have proven effective in domains like math, systems research, and algorithmic discovery, but their

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities

Model ReleasesDGX agent

arXiv:2605.10810v1 Announce Type: new Abstract: We introduce an automatically generated benchmark for predicting hidden text in technical papers. A paper supplies visible context X and a hidden contin

llm 0.32a2

Model ReleasesDGX agent

Release: llm 0.32a2 A bunch of useful stuff in this LLM alpha, but the most important detail is this one: Most reasoning-capable OpenAI models now use the /v1/responses endpoint instead of /v1/chat/co

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

Model ReleasesDGX agent

arXiv:2509.20909v2 Announce Type: replace Abstract: Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memor

MBP-KT: Learning Global Collaborative Information from Meta-Behavioral Pattern for Enhanced Knowledge Tracing

Model ReleasesDGX agent

arXiv:2605.08697v1 Announce Type: new Abstract: The emerging collaborative information-based knowledge tracing (KT) has been a promising way to enhance modeling of learners' knowledge states. The core

MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching

Model ReleasesDGX agent

arXiv:2605.08557v1 Announce Type: cross Abstract: Parameter-efficient adaptation of pretrained vision models is commonly performed through linear probes, prompts, low-rank updates, or lightweight resi

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

Model ReleasesDGX agent

arXiv:2605.10067v1 Announce Type: cross Abstract: Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing ap

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

Model ReleasesDGX agent

arXiv:2605.10616v1 Announce Type: cross Abstract: Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generaliza

NARRA-Gym for Evaluating Interactive Narrative Agents

Model ReleasesDGX agent

arXiv:2605.08503v1 Announce Type: new Abstract: Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmark

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning

ResearchDGX agent

arXiv:2605.08221v1 Announce Type: cross Abstract: This paper presents NoisyCoconut, a novel inference-time method that enhances large language model (LLM) reliability by manipulating internal represen

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

Model ReleasesDGX agent

arXiv:2605.09996v1 Announce Type: new Abstract: While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, wit

Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning

Model ReleasesDGX agent

arXiv:2605.09822v1 Announce Type: cross Abstract: We define Oracle Poisoning, an attack class in which an adversary corrupts a structured knowledge graph that AI agents query at runtime via tool-use p

Phoenix-VL 1.5 Medium Technical Report

Model ReleasesDGX agent

arXiv:2605.10391v1 Announce Type: cross Abstract: We introduce Phoenix-VL 1.5 Medium, a 123B-parameter natively multimodal and multilingual foundation model, adapted to regional languages and the Sing

ProactBench: Beyond What The User Asked For

Model ReleasesDGX agent

arXiv:2605.09228v1 Announce Type: cross Abstract: Most LLM benchmarks score how well a model responds to explicit requests. They leave unmeasured a different conversational ability: noticing and actin

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark

Model ReleasesDGX agent

arXiv:2509.26574v4 Announce Type: replace Abstract: While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason

Reliable LLM-Based Edge-Cloud-Expert Cascades for Telecom Knowledge Systems

Model ReleasesDGX agent

arXiv:2512.20012v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are emerging as key enablers of automation in domains such as telecommunications, assisting with tasks including

ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions

Model ReleasesDGX agent

arXiv:2605.08197v1 Announce Type: cross Abstract: Most causal benchmarks for language models score local answers or graph structure. We introduce ReplaySCM, a 1,300 item benchmark for executable causa

SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing

Model ReleasesDGX agent

arXiv:2605.10831v1 Announce Type: cross Abstract: Large language models possess strong chemical reasoning capabilities, making them effective molecular editors. However, property-relevant information

Split CNN Inference on Networked Microcontrollers

ResearchDGX agent

arXiv:2605.09357v1 Announce Type: cross Abstract: Running deep neural networks on microcontroller units (MCUs) is severely constrained by limited memory resources. While TinyML techniques reduce model

Text-Guided Multi-Scale Frequency Representation Adaptation

Model ReleasesDGX agent

arXiv:2605.08181v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning methods introduce a small number of training parameters, enabling pre-trained models to adapt rapidly to new data dist

The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs

Model ReleasesDGX agent

arXiv:2605.09844v1 Announce Type: new Abstract: The Metacognitive Probe is an exploratory five-task, 15-slot diagnostic that decomposes an LLM's confidence behaviour into five behaviourally-distinct d

The Realignment Problem: When Right becomes Wrong in LLMs

Model ReleasesDGX agent

arXiv:2511.02623v2 Announce Type: replace Abstract: Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over tim

The Silent Vote: Improving Zero-Shot LLM Reliability by Aggregating Semantic Neighborhoods

Model ReleasesDGX agent

arXiv:2605.09739v1 Announce Type: cross Abstract: Large Language Models are increasingly used as zero-shot classifiers in complex reasoning tasks. However, standard constrained decoding suffers from a

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning

Model ReleasesDGX agent

arXiv:2605.09544v1 Announce Type: new Abstract: Tool-integrated reasoning has emerged as a promising paradigm for enhancing large language models with external computation, retrieval, and execution ca

Trajectory Supervision for Continual Tool-Use Learning in LLMs

Model ReleasesDGX agent

arXiv:2605.09734v1 Announce Type: cross Abstract: Most language-model training data shows final artifacts, not the process that produced them. We study a tractable version of this question in tool use

VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning

Model ReleasesDGX agent

arXiv:2604.03701v2 Announce Type: replace Abstract: Video-based numerical reasoning provides a premier arena for testing whether Vision-Language Models (VLMs) truly 'understand' real-world dynamics, a

What Time Is It? How Data Geometry Makes Time Conditioning Optional for Flow Matching

ApplicationsDGX agent

arXiv:2605.08344v1 Announce Type: new Abstract: Recent work has shown that models flow matching models can be trained without explicit time conditioning, challenging the standard view that the interpo

When Less is More: The LLM Scaling Paradox in Context Compression

ResearchDGX agent

arXiv:2602.09789v3 Announce Type: replace Abstract: Scaling up model parameters has long been a prevalent training paradigm driven by the assumption that larger models yield superior generation capabi

Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off

Local AiDGX agent

arXiv:2605.08878v1 Announce Type: cross Abstract: Aligned large language models (LLMs) remain vulnerable to jailbreak attacks. Recent mechanistic studies have identified latent features and representa

ZAYA1-VL-8B Technical Report

ResearchDGX agent

arXiv:2605.08560v1 Announce Type: cross Abstract: We present ZAYA1-VL-8B, a compact mixture-of-experts vision-language model built upon our in-house language model, ZAYA1-8B. Despite its compact size,

11 May 2026

An Embarrassingly Simple Graph Heuristic Reveals Shortcut-Solvable Benchmarks for Sequential Recommendation

Model ReleasesDGX agent

arXiv:2605.07125v1 Announce Type: cross Abstract: Sequential recommendation has increasingly shifted toward generative recommenders that combine sequential patterns with semantic item information. Yet

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?

Model ReleasesDGX agent

arXiv:2605.07937v1 Announce Type: new Abstract: Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreve

Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs

Model ReleasesDGX agent

arXiv:2605.07562v1 Announce Type: new Abstract: Remote sensing vision-language models (RS-VLMs) face a fundamental mismatch with natural-image counterparts: the same geographic object exhibits radical

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation

HardwareDGX agent

arXiv:2605.07985v1 Announce Type: cross Abstract: Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, s

FactoryBench: Evaluating Industrial Machine Understanding

Model ReleasesDGX agent

arXiv:2605.07675v1 Announce Type: new Abstract: We introduce FactoryBench, a benchmark for evaluating time-series models and LLMs on machine understanding over industrial robotic telemetry. Q&A pairs

Fidel-TS: A High-Fidelity Multimodal Benchmark for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2509.24789v4 Announce Type: replace Abstract: The evaluation of time series forecasting models is hindered by a lack of high-quality benchmarks, leading to overestimated assessments of progress.

Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

ResearchDGX agent

arXiv:2602.14868v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in language models. However, relying on sparse

How Value Induction Reshapes LLM Behaviour

SafetyDGX agent

arXiv:2605.07925v1 Announce Type: new Abstract: Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and em

Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs

Model ReleasesDGX agent

arXiv:2605.05957v2 Announce Type: replace Abstract: LLMs reliably correct false claims when presented in isolation, yet when the same claims are embedded in task-oriented requests, they often comply r

MAS-Algorithm: A Workflow for Solving Algorithmic Programming Problems with a Multi-Agent System

Model ReleasesDGX agent

arXiv:2605.05949v2 Announce Type: replace Abstract: Algorithmic problem solving serves as a rigorous testbed for evaluating structured reasoning in AI coding systems, as it directly reflects a model's

MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text

Model ReleasesDGX agent

arXiv:2605.06903v1 Announce Type: cross Abstract: Large language models are now embedded in everyday writing workflows, making reliable AI-generated text detection important for academic integrity, co

On the Invariance and Generality of Neural Scaling Laws

Model ReleasesDGX agent

arXiv:2605.07546v1 Announce Type: new Abstract: Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocatio

Pretraining Induces a Reusable Spectral Basis for Downstream Task Adaptation

Model ReleasesDGX agent

arXiv:2605.07302v1 Announce Type: new Abstract: Finetuning pretrained models occurs in a low-dimensional subspace of the full parameter space. Prior work has focused on characterizing this optimizatio

Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding

Model ReleasesDGX agent

arXiv:2605.07141v1 Announce Type: cross Abstract: Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large lang

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

Model ReleasesDGX agent

arXiv:2605.05995v2 Announce Type: replace-cross Abstract: The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constrain

SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation

Model ReleasesDGX agent

arXiv:2603.05117v3 Announce Type: replace Abstract: Imitation Learning (IL) enables robots to acquire manipulation skills from expert demonstrations. Diffusion Policy (DP) models multi-modal expert be

Synergistic Benefits of Joint Molecule Generation and Property Prediction

ResearchDGX agent

arXiv:2504.16559v3 Announce Type: replace Abstract: Modeling the joint distribution of data samples and their properties allows to construct a single model for both data generation and property predic

Test-Time Compute Games

Model ReleasesDGX agent

arXiv:2601.21839v2 Announce Type: replace-cross Abstract: Test-time compute has emerged as a promising strategy to enhance the reasoning abilities of large language models (LLMs). However, this strate

The Position Curse: LLMs Struggle to Locate the Last Few Items in a List

Model ReleasesDGX agent

arXiv:2605.07127v1 Announce Type: cross Abstract: Modern large language models (LLMs) can find a needle in a haystack (locating a single relevant fact buried among hundreds of thousands of irrelevant

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

Model ReleasesDGX agent

arXiv:2605.06772v1 Announce Type: new Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical questi

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

Model ReleasesDGX agent

arXiv:2605.07114v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language model

Why Self-Inconsistency Arises in GNN Explanations and How to Exploit It

Model ReleasesDGX agent

arXiv:2605.07527v1 Announce Type: cross Abstract: Recent work has observed that explanations produced by Self-Interpretable Graph Neural Networks (SI-GNNs) can be self-inconsistent: when the model is

9 May 2026

ERNIE 5.1 is here 🚀 ERNIE 5.1 significantly reduces pretraining cost while compressing total parameters to ~1/3 and activated parameters to…

Model ReleasesDGX agent

ERNIE 5.1 is here 🚀 ERNIE 5.1 significantly reduces pretraining cost while compressing total parameters to ~1/3 and activated parameters to ~1/2 — using only ~6% of the pretraining cost compared to mo

7 May 2026

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB)

Model ReleasesDGX agent

arXiv:2605.04556v1 Announce Type: cross Abstract: The Massive Sound Embedding Benchmark (MSEB) has emerged as a standard for evaluating the functional breadth of audio models. While initial baselines

← Previous
1…315316317318319…1041
Next →