AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
49,435 results
Local Ai

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

DGX agent

arXiv:2601.15724v2 Announce Type: replace Abstract: Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning

local-aiarxiv-cs-cv
21 Apr 2026
Tutorials
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

When Can LLMs Learn to Reason with Weak Supervision?

DGX agent

arXiv:2604.18574v1 Announce Type: new Abstract: Large language models have achieved significant reasoning improvements through reinforcement learning with verifiable rewards (RLVR). Yet as model capab

tutorialsarxiv-cs-lg
21 Apr 2026
Research

(1D) Ordered Tokens Enable Efficient Test-Time Search

DGX agent

arXiv:2604.15453v1 Announce Type: cross Abstract: Tokenization is a key component of autoregressive (AR) generative models, converting raw data into more manageable units for modeling. Commonly, token

researcharxiv-cs-ai
20 Apr 2026
Model Releases

Aletheia: Gradient-Guided Layer Selection for Efficient LoRA Fine-Tuning Across Architectures

DGX agent

arXiv:2604.15351v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become the dominant parameter-efficient fine-tuning method for large language models, yet standard practice applies LoR

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency

DGX agent

arXiv:2604.16158v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex tasks. Yet ensuring that the reasoning trace both

model-releasesarxiv-cs-ai
20 Apr 2026
Research

Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms

DGX agent

arXiv:2604.15842v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive capabilities, yet their internal mechanisms for handling reasoning-intensive tasks remain unde

researcharxiv-cs-cl
20 Apr 2026
Model Releases

MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition

DGX agent

arXiv:2604.16009v1 Announce Type: new Abstract: Metacognition, the ability to monitor and regulate one's own reasoning, remains under-evaluated in AI benchmarking. We introduce MEDLEY-BENCH, a benchma

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals

DGX agent

arXiv:2604.15859v1 Announce Type: cross Abstract: Forecasting has become a natural benchmark for reasoning under uncertainty. Yet existing evaluations of large language models remain limited to judgme

model-releasesarxiv-cs-ai
20 Apr 2026
Research

Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes

DGX agent

arXiv:2604.14914v1 Announce Type: new Abstract: Text-driven inversion of generative models is a core paradigm for manipulating 2D or 3D content, unlocking numerous applications such as text-based edit

researcharxiv-cs-cv
17 Apr 2026
Research

C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions

DGX agent

arXiv:2604.13521v1 Announce Type: new Abstract: Neural network models with latent recurrent processing, where identical layers are recursively applied to the latent state, have gained attention as pro

researcharxiv-cs-lg
16 Apr 2026
Applications

Diagnostics for Individual-Level Prediction Instability in Machine Learning for Healthcare

DGX agent

arXiv:2603.00192v2 Announce Type: replace Abstract: In healthcare, predictive models increasingly inform patient-level decisions, yet little attention is paid to the variability in individual risk est

applicationsarxiv-cs-lg
16 Apr 2026
Research

Structure- and Stability-Preserving Learning of Port-Hamiltonian Systems

DGX agent

arXiv:2604.13297v1 Announce Type: cross Abstract: This paper investigates the problem of data-driven modeling of port-Hamiltonian systems while preserving their intrinsic Hamiltonian structure and sta

researcharxiv-cs-lg
16 Apr 2026
Research

Accelerating Speculative Decoding with Block Diffusion Draft Trees

DGX agent

arXiv:2604.12989v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive language models by using a lightweight drafter to propose multiple future tokens, which the target model

researcharxiv-cs-cl
15 Apr 2026
Model Releases

Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss

DGX agent

arXiv:2604.12911v1 Announce Type: cross Abstract: Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to p

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents

DGX agent

arXiv:2510.10073v2 Announce Type: replace-cross Abstract: Large vision-language model (LVLM)-based web agents are emerging as powerful tools for automating complex online tasks. However, when deployed

model-releasesarxiv-cs-cv
15 Apr 2026
Research

A Mathematical Explanation of Transformers

DGX agent

arXiv:2510.03989v2 Announce Type: replace-cross Abstract: The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models

researcharxiv-cs-ai
14 Apr 2026
Model Releases

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

DGX agent

arXiv:2601.11044v3 Announce Type: replace Abstract: Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. Howev

model-releasesarxiv-cs-ai
14 Apr 2026
Tutorials

Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation

DGX agent

arXiv:2510.10925v2 Announce Type: replace-cross Abstract: Training student models on synthetic data generated by strong teacher models is a promising way to distilling the capabilities of teachers. Ho

tutorialsarxiv-cs-cl
14 Apr 2026
Model Releases

GIANTS: Generative Insight Anticipation from Scientific Literature

DGX agent

arXiv:2604.09793v1 Announce Type: cross Abstract: Scientific breakthroughs often emerge from synthesizing prior ideas into novel contributions. While language models (LMs) show promise in scientific d

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Incentivizing Honesty among Competitors in Collaborative Learning and Optimization

DGX agent

arXiv:2305.16272v5 Announce Type: replace Abstract: Collaborative learning techniques have the potential to enable training machine learning models that are superior to models trained on a single enti

model-releasesarxiv-cs-lg
14 Apr 2026
Model Releases

Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

DGX agent

arXiv:2604.11666v1 Announce Type: cross Abstract: As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dial

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text

DGX agent

arXiv:2601.17172v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly capable of generating personalized, persuasive text at scale, raising new questions about bias a

model-releasesarxiv-cs-ai
14 Apr 2026
Applications

Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs

DGX agent

arXiv:2604.10495v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in real-world applications, reliable uncertainty quantification (UQ) becomes critical for safe

applicationsarxiv-cs-cl
14 Apr 2026
Model Releases

A Compact Hybrid Convolution--Frequency State Space Network for Learned Image Compression

DGX agent

arXiv:2511.20151v2 Announce Type: replace Abstract: Learned image compression (LIC) has recently benefited from Transformer- and state space models (SSM)- based backbones for modeling long-range depen

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

Adaptive Simulation Experiment for LLM Policy Optimization

DGX agent

arXiv:2604.08779v1 Announce Type: new Abstract: Large language models (LLMs) have significant potential to improve operational efficiency in operations management. Deploying these models requires spec

model-releasesarxiv-cs-lg
13 Apr 2026
Applications

AI Driven Soccer Analysis Using Computer Vision

DGX agent

arXiv:2604.08722v1 Announce Type: cross Abstract: Sport analysis is crucial for team performance since it provides actionable data that can inform coaching decisions, improve player performance, and e

applicationsarxiv-cs-ai
13 Apr 2026
Research

BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation

DGX agent

arXiv:2604.09497v1 Announce Type: cross Abstract: Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases.

researcharxiv-cs-ai
13 Apr 2026
Safety

Bias-Constrained Diffusion Schedules for PDE Emulations: Reconstruction Error Minimization and Efficient Unrolled Training

DGX agent

arXiv:2604.08357v2 Announce Type: replace Abstract: Conditional Diffusion Models are powerful surrogates for emulating complex spatiotemporal dynamics, yet they often fail to match the accuracy of det

safetyarxiv-cs-lg
13 Apr 2026
Model Releases

On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs

DGX agent

arXiv:2509.25214v3 Announce Type: replace-cross Abstract: As increasingly large pre-trained models are released, deploying them on edge devices for privacy-preserving applications requires effective c

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos

DGX agent

arXiv:2604.09037v1 Announce Type: cross Abstract: Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but over

model-releasesarxiv-cs-cl
13 Apr 2026
Applications

Task-Distributionally Robust Data-Free Meta-Learning

DGX agent

arXiv:2311.14756v2 Announce Type: replace-cross Abstract: Data-Free Meta-Learning (DFML) aims to enable efficient learning of unseen few-shot tasks, by meta-learning from multiple pre-trained models w

applicationsarxiv-cs-ai
13 Apr 2026
Model Releases

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection

DGX agent

arXiv:2511.20162v2 Announce Type: replace Abstract: Large multi-modal models (LMMs) show increasing performance in realistic visual tasks for images and, more recently, for videos. For example, given

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting

DGX agent

arXiv:2601.02670v2 Announce Type: replace Abstract: We introduce self-jailbreaking, a threat model in which an aligned LLM guides its own compromise. Unlike most jailbreak techniques, which oft

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Calibration of a neural network ocean closure for improved mean state and variability

DGX agent

arXiv:2604.06398v1 Announce Type: cross Abstract: Global ocean models exhibit biases in the mean state and variability, particularly at coarse resolution, where mesoscale eddies are unresolved. To add

model-releasesarxiv-cs-lg
10 Apr 2026
Model Releases

Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment

DGX agent

arXiv:2601.08258v3 Announce Type: replace Abstract: Large language models increasingly fail in a way that scalar accuracy cannot diagnose: they produce a sound reasoning trace and then abandon it unde

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Dynamic Context Evolution for Scalable Synthetic Data Generation

DGX agent

arXiv:2604.07147v1 Announce Type: cross Abstract: Large language models produce repetitive output when prompted independently across many batches, a phenomenon we term cross-batch mode collapse: the p

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency

DGX agent

arXiv:2604.03044v2 Announce Type: replace-cross Abstract: We introduce JoyAI-LLM Flash, an efficient Mixture-of-Experts (MoE) language model designed to redefine the trade-off between strong performan

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces

DGX agent

arXiv:2604.06188v1 Announce Type: cross Abstract: People increasingly hold sustained, open-ended conversations with large language models (LLMs). Public reports and early studies suggest that, in such

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

LoFT: Parameter-Efficient Fine-Tuning for Long-tailed Semi-Supervised Learning in Open-World Scenarios

DGX agent

arXiv:2509.09926v5 Announce Type: replace Abstract: Long-tailed semi-supervised learning (LTSSL) presents a formidable challenge where models must overcome the scarcity of tail samples while mitigatin

model-releasesarxiv-cs-lg
10 Apr 2026
Model Releases

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

DGX agent

arXiv:2604.04771v2 Announce Type: replace-cross Abstract: Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remain

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

PLUME: Latent Reasoning Based Universal Multimodal Embedding

DGX agent

arXiv:2604.02073v2 Announce Type: replace Abstract: Universal multimodal embedding (UME) maps heterogeneous inputs into a shared retrieval space with a single model. Recent approaches improve UME by g

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards

DGX agent

arXiv:2512.01236v2 Announce Type: replace Abstract: Personalized generation models for a single subject have demonstrated remarkable effectiveness, highlighting their significant potential. However, w

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition

DGX agent

arXiv:2604.07884v1 Announce Type: new Abstract: High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory an

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

SALLIE: Safeguarding Against Latent Language & Image Exploits

DGX agent

arXiv:2604.06247v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) remain highly vulnerable to textual and visual jailbreaks, as well as prompt injections

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

T-Gated Adapter: A Lightweight Temporal Adapter for Vision-Language Medical Segmentation

DGX agent

arXiv:2604.08167v1 Announce Type: new Abstract: Medical image segmentation traditionally relies on fully supervised 3D architectures that demand a large amount of dense, voxel-level annotations from c

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

TEMPER: Testing Emotional Perturbation in Quantitative Reasoning

DGX agent

arXiv:2604.07801v1 Announce Type: new Abstract: Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world quer

model-releasesarxiv-cs-cl
10 Apr 2026
Research

Federated Compositional Muon Optimizer for Matrix-Wise Models

DGX agent

arXiv:2608.12710v1 Announce Type: new Abstract: Muon, a more recently developed optimizer, is useful for matrix-wise models in AI areas. Although many works have studied Muon and its variants, these m

researcharxiv-cs-lg
14 Aug 2026
Safety

Large Language Models Persuade Without Planning Theory of Mind

DGX agent

arXiv:2602.17045v2 Announce Type: replace Abstract: A growing body of work attempts to evaluate the theory of mind (ToM) abilities of humans and large language models (LLMs) using static, non-interact

safetyarxiv-cs-cl
14 Aug 2026
← Previous
1…200201202203204…1030
Next →