AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,106 results
Model Releases

Scaling Laws and Tradeoffs in Recurrent Networks of Expressive Neurons

DGX agent

arXiv:2605.12049v1 Announce Type: new Abstract: Cortical neurons are complex, multi-timescale processors wired into recurrent circuits, shaped by long evolutionary pressure under stringent biological

model-releasesarxiv-cs-lg
13 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Search Your Block Floating Point Scales!

DGX agent

arXiv:2605.12464v1 Announce Type: new Abstract: Quantization has emerged as a standard technique for accelerating inference for generative models by enabling faster low-precision computations and redu

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Self-Supervised Laplace Approximation for Bayesian Uncertainty Quantification

DGX agent

arXiv:2605.12208v1 Announce Type: cross Abstract: Approximate Bayesian inference typically revolves around computing the posterior parameter distribution. In practice, however, the main object of inte

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens

DGX agent

arXiv:2602.15620v4 Announce Type: replace Abstract: Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning

DGX agent

arXiv:2605.11922v1 Announce Type: cross Abstract: Existing code reasoning methods primarily supervise final code outputs, ignoring intermediate states, often leading to reward hacking where correct an

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization

DGX agent

arXiv:2605.08704v1 Announce Type: new Abstract: Multi-agent reasoning has shown promise for improving the problem-solving ability of large language models by allowing multiple agents to explore divers

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design

DGX agent

arXiv:2605.08756v1 Announce Type: new Abstract: Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent works show t

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

DGX agent

arXiv:2601.01762v2 Announce Type: replace-cross Abstract: Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. W

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment

DGX agent

arXiv:2603.26680v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) evolve into lifelong AI assistants, LLM personalization has become a critical frontier. However, progress is c

model-releasesarxiv-cs-ai
12 May 2026
Research

An Elastic Shape Variational Autoencoder for Skeleton Pose Trajectories

DGX agent

arXiv:2605.09231v1 Announce Type: new Abstract: Deep generative models provide flexible frameworks for modeling complex, structured data such as images, videos, 3D objects, and texts. However, when ap

researcharxiv-cs-cv
12 May 2026
Model Releases

Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness

DGX agent

arXiv:2605.09634v1 Announce Type: new Abstract: LLMs can estimate Hospital Anxiety and Depression Scale (HADS) scores from speech in a zero-shot manner, but clinical deployment requires reliability ac

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts

DGX agent

arXiv:2503.05066v5 Announce Type: replace-cross Abstract: The Mixture of Experts (MoE) is an effective architecture for scaling large language models by leveraging sparse expert activation to balance

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure

DGX agent

arXiv:2605.08740v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) decompose transformer residual streams into interpretable feature dictionaries, yet the relationship between SAE width and

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

ChartDiff: A Large-Scale Benchmark for Comprehending Pairs of Charts

DGX agent

arXiv:2603.28902v2 Announce Type: replace Abstract: Charts are central to analytical reasoning, yet existing benchmarks for chart understanding focus almost exclusively on single-chart interpretation

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics

DGX agent

arXiv:2605.09584v1 Announce Type: cross Abstract: Inpatient clinical reasoning is a sequential decision under partial observability: the clinician sees the admission so far and must choose the next ac

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization

DGX agent

arXiv:2605.08802v1 Announce Type: new Abstract: Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perfo

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation

DGX agent

arXiv:2605.08522v1 Announce Type: new Abstract: The evaluation of Large Language Models (LLMs) faces a critical challenge in construct validity, where fragmented benchmarks and ad hoc metrics frequent

model-releasesarxiv-cs-cl
12 May 2026
Research

Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents

DGX agent

arXiv:2605.08442v1 Announce Type: cross Abstract: Persistent memory attacks against LLM agents achieve high attack success rates against open-source models. In these attacks, malicious instructions in

researcharxiv-cs-ai
12 May 2026
Model Releases

DiffATS: Diffusion in Aligned Tensor Space

DGX agent

arXiv:2605.09275v1 Announce Type: new Abstract: Direct diffusion modeling of high-resolution spatiotemporal fields is computationally challenging. Parameter-efficient primitives address this by repres

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Different Prompts, Different Ranks: Prompt-aware Dynamic Rank Selection for SVD-based LLM Compression

DGX agent

arXiv:2605.08568v1 Announce Type: new Abstract: Large language models (LLMs) have rapidly grown in scale, creating substantial memory and computational costs that hinder efficient deployment. Singular

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments

DGX agent

arXiv:2503.06047v2 Announce Type: replace Abstract: Large language model (LLM)-based agents are increasingly applied to complex strategic environments that demand long-horizon reasoning, multi-agent i

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

EdgeFlowerTune: Evaluating Federated LLM Fine-Tuning Under Realistic Edge System Constraints

DGX agent

arXiv:2605.08636v1 Announce Type: new Abstract: Federated fine-tuning offers a promising paradigm for adapting large language models (LLMs) on edge devices by leveraging the rich, diverse, and continu

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Efficient Evaluation of LLM Performance with Statistical Guarantees

DGX agent

arXiv:2601.20251v3 Announce Type: replace-cross Abstract: Exhaustively evaluating many large language models (LLMs) on a large suite of benchmarks is expensive. We cast benchmarking as finite-populati

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Efficient Neural Architectures for Real-Time ECG Interpretation on Limited Hardware

DGX agent

arXiv:2605.09848v1 Announce Type: new Abstract: Electrocardiogram (ECG) interpretation is essential for diagnosing a wide range of cardiac abnormalities. While deep learning has shown strong potential

model-releasesarxiv-cs-lg
12 May 2026
Research

EMO: Pretraining Mixture of Experts for Emergent Modularity

DGX agent

arXiv:2605.06663v2 Announce Type: replace Abstract: Large language models are typically deployed as monolithic systems, requiring the full model even when applications need only a narrow subset of cap

researcharxiv-cs-cl
12 May 2026
Model Releases

Fashion Florence: Fine-Tuning Florence-2 for Structured Fashion Attribute Extraction

DGX agent

arXiv:2605.09827v1 Announce Type: cross Abstract: We present Fashion Florence, a Florence-2 vision-language model fine-tuned with LoRA to extract structured fashion attributes from clothing images. Gi

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Feature Rivalry in Sparse Autoencoder Representations: A Mechanistic Study of Uncertainty-Driven Feature Competition in LLMs

DGX agent

arXiv:2605.08149v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) decompose large language model representations into interpretable features, but how these features interact under uncertain

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain

DGX agent

arXiv:2605.09106v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in financial contexts, raising critical concerns about reliability, alignment, and susceptibility

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

FRACTAL: SSM with Fractional Recurrent Architecture for Computational Temporal Analysis of Long Sequences

DGX agent

arXiv:2605.08833v1 Announce Type: new Abstract: Effective sequence modeling fundamentally requires balancing the retention of unbounded history with the high-resolution detection of abrupt short-term

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Generating Leakage-Free Benchmarks for Robust RAG Evaluation

DGX agent

arXiv:2605.08838v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is widely used to augment large language models (LLMs) with external knowledge. However, many benchmark datasets,

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Geometry-free prediction of inertial lift forces in microfluidic devices using deep learning

DGX agent

arXiv:2605.08109v1 Announce Type: new Abstract: Inertial microfluidic devices (IMDs) offer low-cost, high-throughput alternative techniques for many traditional particle- (or cell-) manipulation tasks

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking

DGX agent

arXiv:2605.10893v1 Announce Type: new Abstract: Large vision-language models suffer from visual ungroundedness: they can produce a fluent, confident, and even correct response driven entirely by langu

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

HS-FNO: History-Space Fourier Neural Operator for Non-Markovian Partial Differential Equations

DGX agent

arXiv:2605.09523v1 Announce Type: new Abstract: Neural operators provide fast surrogate models for time-dependent partial differential equations, but their standard autoregressive use usually assumes

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

In-Context Fixation: When Demonstrated Labels Override Semantics in Few-Shot Classification

DGX agent

arXiv:2605.08295v1 Announce Type: cross Abstract: While random demonstration labels barely hurt in-context learning (Min et al., 2022), we show that homogeneous labels--even semantically valid ones--c

model-releasesarxiv-cs-ai
12 May 2026
Applications

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

DGX agent

arXiv:2601.22143v2 Announce Type: replace-cross Abstract: Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented abilit

applicationsarxiv-cs-cv
12 May 2026
Model Releases

LEVI: Stronger Search Architectures Can Substitute for Larger LLMs in Evolutionary Search

DGX agent

arXiv:2605.09764v1 Announce Type: cross Abstract: LLM-guided evolutionary methods such as AlphaEvolve have proven effective in domains like math, systems research, and algorithmic discovery, but their

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities

DGX agent

arXiv:2605.10810v1 Announce Type: new Abstract: We introduce an automatically generated benchmark for predicting hidden text in technical papers. A paper supplies visible context X and a hidden contin

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

DGX agent

arXiv:2509.20909v2 Announce Type: replace Abstract: Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memor

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

MBP-KT: Learning Global Collaborative Information from Meta-Behavioral Pattern for Enhanced Knowledge Tracing

DGX agent

arXiv:2605.08697v1 Announce Type: new Abstract: The emerging collaborative information-based knowledge tracing (KT) has been a promising way to enhance modeling of learners' knowledge states. The core

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching

DGX agent

arXiv:2605.08557v1 Announce Type: cross Abstract: Parameter-efficient adaptation of pretrained vision models is commonly performed through linear probes, prompts, low-rank updates, or lightweight resi

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

DGX agent

arXiv:2605.10067v1 Announce Type: cross Abstract: Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing ap

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

DGX agent

arXiv:2605.10616v1 Announce Type: cross Abstract: Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generaliza

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

NARRA-Gym for Evaluating Interactive Narrative Agents

DGX agent

arXiv:2605.08503v1 Announce Type: new Abstract: Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmark

model-releasesarxiv-cs-cl
12 May 2026
Research

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning

DGX agent

arXiv:2605.08221v1 Announce Type: cross Abstract: This paper presents NoisyCoconut, a novel inference-time method that enhances large language model (LLM) reliability by manipulating internal represen

researcharxiv-cs-ai
12 May 2026
Model Releases

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

DGX agent

arXiv:2605.09996v1 Announce Type: new Abstract: While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, wit

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning

DGX agent

arXiv:2605.09822v1 Announce Type: cross Abstract: We define Oracle Poisoning, an attack class in which an adversary corrupts a structured knowledge graph that AI agents query at runtime via tool-use p

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Phoenix-VL 1.5 Medium Technical Report

DGX agent

arXiv:2605.10391v1 Announce Type: cross Abstract: We introduce Phoenix-VL 1.5 Medium, a 123B-parameter natively multimodal and multilingual foundation model, adapted to regional languages and the Sing

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

ProactBench: Beyond What The User Asked For

DGX agent

arXiv:2605.09228v1 Announce Type: cross Abstract: Most LLM benchmarks score how well a model responds to explicit requests. They leave unmeasured a different conversational ability: noticing and actin

model-releasesarxiv-cs-ai
12 May 2026
← Previous
1…324325326327328…1065
Next →