AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,106 results
Model Releases

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark

DGX agent

arXiv:2509.26574v4 Announce Type: replace Abstract: While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason

model-releasesarxiv-cs-ai
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Reliable LLM-Based Edge-Cloud-Expert Cascades for Telecom Knowledge Systems

DGX agent

arXiv:2512.20012v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are emerging as key enablers of automation in domains such as telecommunications, assisting with tasks including

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions

DGX agent

arXiv:2605.08197v1 Announce Type: cross Abstract: Most causal benchmarks for language models score local answers or graph structure. We introduce ReplaySCM, a 1,300 item benchmark for executable causa

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing

DGX agent

arXiv:2605.10831v1 Announce Type: cross Abstract: Large language models possess strong chemical reasoning capabilities, making them effective molecular editors. However, property-relevant information

model-releasesarxiv-cs-ai
12 May 2026
Research

Split CNN Inference on Networked Microcontrollers

DGX agent

arXiv:2605.09357v1 Announce Type: cross Abstract: Running deep neural networks on microcontroller units (MCUs) is severely constrained by limited memory resources. While TinyML techniques reduce model

researcharxiv-cs-lg
12 May 2026
Model Releases

Text-Guided Multi-Scale Frequency Representation Adaptation

DGX agent

arXiv:2605.08181v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning methods introduce a small number of training parameters, enabling pre-trained models to adapt rapidly to new data dist

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs

DGX agent

arXiv:2605.09844v1 Announce Type: new Abstract: The Metacognitive Probe is an exploratory five-task, 15-slot diagnostic that decomposes an LLM's confidence behaviour into five behaviourally-distinct d

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Realignment Problem: When Right becomes Wrong in LLMs

DGX agent

arXiv:2511.02623v2 Announce Type: replace Abstract: Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over tim

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

The Silent Vote: Improving Zero-Shot LLM Reliability by Aggregating Semantic Neighborhoods

DGX agent

arXiv:2605.09739v1 Announce Type: cross Abstract: Large Language Models are increasingly used as zero-shot classifiers in complex reasoning tasks. However, standard constrained decoding suffers from a

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning

DGX agent

arXiv:2605.09544v1 Announce Type: new Abstract: Tool-integrated reasoning has emerged as a promising paradigm for enhancing large language models with external computation, retrieval, and execution ca

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Trajectory Supervision for Continual Tool-Use Learning in LLMs

DGX agent

arXiv:2605.09734v1 Announce Type: cross Abstract: Most language-model training data shows final artifacts, not the process that produced them. We study a tractable version of this question in tool use

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning

DGX agent

arXiv:2604.03701v2 Announce Type: replace Abstract: Video-based numerical reasoning provides a premier arena for testing whether Vision-Language Models (VLMs) truly 'understand' real-world dynamics, a

model-releasesarxiv-cs-cv
12 May 2026
Applications

What Time Is It? How Data Geometry Makes Time Conditioning Optional for Flow Matching

DGX agent

arXiv:2605.08344v1 Announce Type: new Abstract: Recent work has shown that models flow matching models can be trained without explicit time conditioning, challenging the standard view that the interpo

applicationsarxiv-cs-lg
12 May 2026
Research

When Less is More: The LLM Scaling Paradox in Context Compression

DGX agent

arXiv:2602.09789v3 Announce Type: replace Abstract: Scaling up model parameters has long been a prevalent training paradigm driven by the assumption that larger models yield superior generation capabi

researcharxiv-cs-lg
12 May 2026
Local Ai

Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off

DGX agent

arXiv:2605.08878v1 Announce Type: cross Abstract: Aligned large language models (LLMs) remain vulnerable to jailbreak attacks. Recent mechanistic studies have identified latent features and representa

local-aiarxiv-cs-ai
12 May 2026
Research

ZAYA1-VL-8B Technical Report

DGX agent

arXiv:2605.08560v1 Announce Type: cross Abstract: We present ZAYA1-VL-8B, a compact mixture-of-experts vision-language model built upon our in-house language model, ZAYA1-8B. Despite its compact size,

researcharxiv-cs-ai
12 May 2026
Model Releases

An Embarrassingly Simple Graph Heuristic Reveals Shortcut-Solvable Benchmarks for Sequential Recommendation

DGX agent

arXiv:2605.07125v1 Announce Type: cross Abstract: Sequential recommendation has increasingly shifted toward generative recommenders that combine sequential patterns with semantic item information. Yet

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?

DGX agent

arXiv:2605.07937v1 Announce Type: new Abstract: Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreve

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs

DGX agent

arXiv:2605.07562v1 Announce Type: new Abstract: Remote sensing vision-language models (RS-VLMs) face a fundamental mismatch with natural-image counterparts: the same geographic object exhibits radical

model-releasesarxiv-cs-cv
11 May 2026
Hardware

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation

DGX agent

arXiv:2605.07985v1 Announce Type: cross Abstract: Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, s

hardwarearxiv-cs-ai
11 May 2026
Model Releases

FactoryBench: Evaluating Industrial Machine Understanding

DGX agent

arXiv:2605.07675v1 Announce Type: new Abstract: We introduce FactoryBench, a benchmark for evaluating time-series models and LLMs on machine understanding over industrial robotic telemetry. Q&A pairs

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Fidel-TS: A High-Fidelity Multimodal Benchmark for Time Series Forecasting

DGX agent

arXiv:2509.24789v4 Announce Type: replace Abstract: The evaluation of time series forecasting models is hindered by a lack of high-quality benchmarks, leading to overestimated assessments of progress.

model-releasesarxiv-cs-lg
11 May 2026
Research

Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

DGX agent

arXiv:2602.14868v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in language models. However, relying on sparse

researcharxiv-cs-ai
11 May 2026
Safety

How Value Induction Reshapes LLM Behaviour

DGX agent

arXiv:2605.07925v1 Announce Type: new Abstract: Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and em

safetyarxiv-cs-cl
11 May 2026
Model Releases

Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs

DGX agent

arXiv:2605.05957v2 Announce Type: replace Abstract: LLMs reliably correct false claims when presented in isolation, yet when the same claims are embedded in task-oriented requests, they often comply r

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

MAS-Algorithm: A Workflow for Solving Algorithmic Programming Problems with a Multi-Agent System

DGX agent

arXiv:2605.05949v2 Announce Type: replace Abstract: Algorithmic problem solving serves as a rigorous testbed for evaluating structured reasoning in AI coding systems, as it directly reflects a model's

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text

DGX agent

arXiv:2605.06903v1 Announce Type: cross Abstract: Large language models are now embedded in everyday writing workflows, making reliable AI-generated text detection important for academic integrity, co

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

On the Invariance and Generality of Neural Scaling Laws

DGX agent

arXiv:2605.07546v1 Announce Type: new Abstract: Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocatio

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Pretraining Induces a Reusable Spectral Basis for Downstream Task Adaptation

DGX agent

arXiv:2605.07302v1 Announce Type: new Abstract: Finetuning pretrained models occurs in a low-dimensional subspace of the full parameter space. Prior work has focused on characterizing this optimizatio

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding

DGX agent

arXiv:2605.07141v1 Announce Type: cross Abstract: Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large lang

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

DGX agent

arXiv:2605.05995v2 Announce Type: replace-cross Abstract: The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constrain

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation

DGX agent

arXiv:2603.05117v3 Announce Type: replace Abstract: Imitation Learning (IL) enables robots to acquire manipulation skills from expert demonstrations. Diffusion Policy (DP) models multi-modal expert be

model-releasesarxiv-cs-ro
11 May 2026
Research

Synergistic Benefits of Joint Molecule Generation and Property Prediction

DGX agent

arXiv:2504.16559v3 Announce Type: replace Abstract: Modeling the joint distribution of data samples and their properties allows to construct a single model for both data generation and property predic

researcharxiv-cs-lg
11 May 2026
Model Releases

Test-Time Compute Games

DGX agent

arXiv:2601.21839v2 Announce Type: replace-cross Abstract: Test-time compute has emerged as a promising strategy to enhance the reasoning abilities of large language models (LLMs). However, this strate

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

The Position Curse: LLMs Struggle to Locate the Last Few Items in a List

DGX agent

arXiv:2605.07127v1 Announce Type: cross Abstract: Modern large language models (LLMs) can find a needle in a haystack (locating a single relevant fact buried among hundreds of thousands of irrelevant

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

DGX agent

arXiv:2605.06772v1 Announce Type: new Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical questi

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

DGX agent

arXiv:2605.07114v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language model

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Why Self-Inconsistency Arises in GNN Explanations and How to Exploit It

DGX agent

arXiv:2605.07527v1 Announce Type: cross Abstract: Recent work has observed that explanations produced by Self-Interpretable Graph Neural Networks (SI-GNNs) can be self-inconsistent: when the model is

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB)

DGX agent

arXiv:2605.04556v1 Announce Type: cross Abstract: The Massive Sound Embedding Benchmark (MSEB) has emerged as a standard for evaluating the functional breadth of audio models. While initial baselines

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Benchmarking POS Tagging for the Tajik Language: A Comparative Study of Neural Architectures on the TajPersParallel Corpus

DGX agent

arXiv:2605.04576v1 Announce Type: new Abstract: This paper presents the first benchmark for the task of automatic part-of-speech (POS) tagging for the Tajik language. Despite the existence of multilin

model-releasesarxiv-cs-cl
7 May 2026
Hardware

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs

DGX agent

arXiv:2605.04357v1 Announce Type: cross Abstract: The usage of large language models (LLMs) has grown increasingly fragmented, with no single model dominating. Meanwhile, cloud providers offer a wide

hardwarearxiv-cs-cl
7 May 2026
Model Releases

Enhancing Agent Safety Judgment: Controlled Benchmark Rewriting and Analogical Reasoning for Deceptive Out-of-Distribution Scenarios

DGX agent

arXiv:2605.03242v1 Announce Type: new Abstract: Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environm

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression

DGX agent

arXiv:2605.04084v1 Announce Type: new Abstract: Compressing large language models (LLMs) for deployment on commodity GPUs remains challenging: conventional scalar quantization is limited to fixed bit-

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

FlatASCEND: Autoregressive Clinical Sequence Generation with Continuous Time Prediction and Association-Based Pharmacological Testing

DGX agent

arXiv:2605.04071v1 Announce Type: new Abstract: Autoregressive models can predict clinical events, but generating patient-conditioned multi-step trajectories that respond to intervention tokens and te

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs

DGX agent

arXiv:2605.04065v1 Announce Type: new Abstract: Unsupervised reinforcement learning (RL) has emerged as a promising paradigm for enabling self-improvement in large language models (LLMs). However, exi

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Nsanku: Evaluating Zero-Shot Translation Performance of LLMs for Ghanaian Languages

DGX agent

arXiv:2605.04208v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive multilingual capabilities for well-resourced languages, yet their performance on low-resource

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval

DGX agent

arXiv:2605.03361v2 Announce Type: new Abstract: As multimodal content continues to expand at a rapid pace, audio retrieval has emerged as a key enabling technology for media search, content organizati

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

Tree-Conditioned Edit Flows for Ancestral Sequence Reconstruction

DGX agent

arXiv:2605.04119v1 Announce Type: cross Abstract: Ancestral sequence reconstruction (ASR) aims to infer extinct protein sequences at internal nodes of a phylogenetic tree. Classical ASR methods are ty

model-releasesarxiv-cs-lg
7 May 2026
← Previous
1…325326327328329…1065
Next →