AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
49,435 results
Model Releases

METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models

DGX agent

arXiv:2604.11502v1 Announce Type: cross Abstract: Contextual causal reasoning is a critical yet challenging capability for Large Language Models (LLMs). Existing benchmarks, however, often evaluate th

model-releasesarxiv-cs-ai
14 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models

DGX agent

arXiv:2604.10971v1 Announce Type: cross Abstract: In the progress of industrial anomaly detection, general anomaly detection (GAD) is an emerging trend and also the ultimate goal. Unlike the conventio

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models

DGX agent

arXiv:2604.10866v1 Announce Type: new Abstract: AI agents are expected to perform professional work across hundreds of occupational domains (from emergency department triage to nuclear reactor safety

model-releasesarxiv-cs-cl
14 Apr 2026
Research

Physics and causally constrained discrete-time neural models of turbulent dynamical systems

DGX agent

arXiv:2602.13847v3 Announce Type: replace-cross Abstract: We present a framework for constructing physics and causally constrained neural models of turbulent dynamical systems from data. We first form

researcharxiv-cs-lg
14 Apr 2026
Research

PoreDiT: A Scalable Generative Model for Large-Scale Digital Rock Reconstruction

DGX agent

arXiv:2604.10171v1 Announce Type: new Abstract: This manuscript presents PoreDiT, a novel generative model designed for high-efficiency digital rock reconstruction at gigavoxel scales. Addressing the

researcharxiv-cs-ai
14 Apr 2026
Research

Resource Consumption Threats in Large Language Models

DGX agent

arXiv:2603.16068v3 Announce Type: replace-cross Abstract: Given limited and costly computational infrastructure, resource efficiency is a key requirement for large language models (LLMs). Efficient LL

researcharxiv-cs-ai
14 Apr 2026
Research

Rethinking the Diffusion Model from a Langevin Perspective

DGX agent

arXiv:2604.10465v1 Announce Type: cross Abstract: Diffusion models are often introduced from multiple perspectives, such as VAEs, score matching, or flow matching, accompanied by dense and technically

researcharxiv-cs-ai
14 Apr 2026
Model Releases

SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors

DGX agent

arXiv:2510.17516v4 Announce Type: replace-cross Abstract: Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only i

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models

DGX agent

arXiv:2604.10733v1 Announce Type: cross Abstract: Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

TraversalBench: Challenging Paths to Follow for Vision Language Models

DGX agent

arXiv:2604.10999v1 Announce Type: new Abstract: Vision-language models (VLMs) perform strongly on many multimodal benchmarks. However, the ability to follow complex visual paths -- a task that human o

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

Vision-Language-Action Model, Robustness, Multi-modal Learning, Robot Manipulation

DGX agent

arXiv:2604.10055v1 Announce Type: new Abstract: Despite their strong performance in embodied tasks, recent Vision-Language-Action (VLA) models remain highly fragile under multimodal perturbations, whe

model-releasesarxiv-cs-ro
14 Apr 2026
Model Releases

Your Model Diversity, Not Method, Determines Reasoning Strategy

DGX agent

arXiv:2604.10827v1 Announce Type: new Abstract: Compute scaling for LLM reasoning requires allocating budget between exploring solution approaches (breadth) and refining promising solutions (depth). M

model-releasesarxiv-cs-ai
14 Apr 2026
Research

Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts

DGX agent

arXiv:2604.09364v1 Announce Type: cross Abstract: When a Vision-Language Model (VLM) sees a blue banana and answers 'yellow', is the problem of perception or arbitration? We explore the question in te

researcharxiv-cs-cl
13 Apr 2026
Model Releases

CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space

DGX agent

arXiv:2604.09029v1 Announce Type: cross Abstract: Large language models have been widely explored as decision-support tools in high-stakes domains due to their contextual understanding and reasoning c

model-releasesarxiv-cs-ai
13 Apr 2026
Safety

From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales

DGX agent

arXiv:2604.08591v1 Announce Type: cross Abstract: Hallucinations in large ASR models present a critical safety risk. In this work, we propose the extit{Spectral Sensitivity Theorem}, which predicts a

safetyarxiv-cs-ai
13 Apr 2026
Model Releases

From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

DGX agent

arXiv:2604.09459v1 Announce Type: new Abstract: Reinforcement learning (RL) for large language models (LLMs) increasingly relies on sparse, outcome-level rewards -- yet determining which actions withi

model-releasesarxiv-cs-cl
13 Apr 2026
Model Releases

Medical Reasoning with Large Language Models: A Survey and MR-Bench

DGX agent

arXiv:2604.08559v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved strong performance on medical exam-style tasks, motivating growing interest in their deployment in real-wor

model-releasesarxiv-cs-ai
13 Apr 2026
Research

No Single Best Model for Diversity: Learning a Router for Sample Diversity

DGX agent

arXiv:2604.02319v2 Announce Type: replace Abstract: When posed with prompts that permit a large number of valid answers, comprehensively generating them is the first step towards satisfying a wide ran

researcharxiv-cs-cl
13 Apr 2026
Model Releases

RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models

DGX agent

arXiv:2511.19704v2 Announce Type: replace Abstract: Open-vocabulary semantic segmentation (OVSS) underpins many vision and robotics tasks that require generalizable semantic understanding. Existing ap

model-releasesarxiv-cs-cv
13 Apr 2026
Tutorials

SkillFactory: Self-Distillation For Learning Cognitive Behaviors

DGX agent

arXiv:2512.04072v2 Announce Type: replace-cross Abstract: Reasoning models leveraging long chains of thought employ various cognitive skills, such as verification of their answers, backtracking, retry

tutorialsarxiv-cs-ai
13 Apr 2026
Safety

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

DGX agent

arXiv:2604.08815v1 Announce Type: new Abstract: Medical vision-language models (VLMs) show strong performance on radiology tasks but often produce fluent yet weakly grounded conclusions due to over-re

safetyarxiv-cs-cv
13 Apr 2026
Model Releases

Beyond Mamba: Enhancing State-space Models with Deformable Dilated Convolutions for Multi-scale Traffic Object Detection

DGX agent

arXiv:2604.08038v1 Announce Type: new Abstract: In a real-world traffic scenario, varying-scale objects are usually distributed in a cluttered background, which poses great challenges to accurate dete

model-releasesarxiv-cs-cv
10 Apr 2026
Research

Code Sharing In Prediction Model Research: A Scoping Review

DGX agent

arXiv:2604.06212v1 Announce Type: cross Abstract: Analytical code is essential for reproducing diagnostic and prognostic prediction model research, yet code availability in the published literature re

researcharxiv-cs-ai
10 Apr 2026
Research

Controller Design for Structured State-space Models via Contraction Theory

DGX agent

arXiv:2604.07069v1 Announce Type: cross Abstract: This paper presents an indirect data-driven output feedback controller synthesis for nonlinear systems, leveraging Structured State-space Models (SSMs

researcharxiv-cs-lg
10 Apr 2026
Tutorials

DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models

DGX agent

arXiv:2306.14685v5 Announce Type: replace-cross Abstract: We demonstrate that pre-trained text-to-image diffusion models, despite being trained on raster images, possess a remarkable capacity to guide

tutorialsarxiv-cs-ai
10 Apr 2026
Model Releases

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing

DGX agent

arXiv:2509.01986v4 Announce Type: replace-cross Abstract: In recent years, integrating multimodal understanding and generation into a single unified model has emerged as a promising paradigm. While th

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task

DGX agent

arXiv:2510.14509v3 Announce Type: replace-cross Abstract: The rapid advancement in large language models (LLMs) has demonstrated significant potential in End-to-End Software Development (E2ESD). Howev

model-releasesarxiv-cs-cl
10 Apr 2026
Local Ai

From Synthetic Data to Real Restorations: Diffusion Model for Patient-specific Dental Crown Completion

DGX agent

arXiv:2603.26588v2 Announce Type: replace-cross Abstract: We present ToothCraft, a diffusion-based model for the contextual generation of tooth crowns, trained on artificially created incomplete teeth

local-aiarxiv-cs-lg
10 Apr 2026
Safety

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs

DGX agent

arXiv:2604.07655v1 Announce Type: cross Abstract: Hard-gated safety checkers often over-refuse and misalign with a vendor's model spec; prevailing taxonomies also neglect robustness and honesty, yield

safetyarxiv-cs-cl
10 Apr 2026
Hardware

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

DGX agent

arXiv:2604.07812v1 Announce Type: new Abstract: In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making th

hardwarearxiv-cs-cv
10 Apr 2026
Safety

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles

DGX agent

arXiv:2604.07650v1 Announce Type: cross Abstract: The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Shared pretra

safetyarxiv-cs-cl
10 Apr 2026
Research

HST-HGN: Heterogeneous Spatial-Temporal Hypergraph Networks with Bidirectional State Space Models for Global Fatigue Assessment

DGX agent

arXiv:2604.08435v1 Announce Type: new Abstract: It remains challenging to assess driver fatigue from untrimmed videos under constrained computational budgets, due to the difficulty of modeling long-ra

researcharxiv-cs-cv
10 Apr 2026
Research

Interventional Time Series Priors for Causal Foundation Models

DGX agent

arXiv:2603.11090v2 Announce Type: replace Abstract: Prior-data fitted networks (PFNs) have emerged as powerful foundation models for tabular causal inference, yet their extension to time series remain

researcharxiv-cs-lg
10 Apr 2026
Model Releases

Negative Binomial Variational Autoencoders for Overdispersed Latent Modeling

DGX agent

arXiv:2508.05423v2 Announce Type: replace Abstract: Although artificial neural networks are often described as brain-inspired, their representations typically rely on continuous activations, such as t

model-releasesarxiv-cs-lg
10 Apr 2026
Model Releases

Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models

DGX agent

arXiv:2604.01840v2 Announce Type: replace Abstract: While Reinforcement Learning from Verifiable Rewards (RLVR) has advanced reasoning in Large Vision-Language Models (LVLMs), prevailing frameworks su

model-releasesarxiv-cs-ai
10 Apr 2026
Research

Phantasia: Context-Adaptive Backdoors in Vision Language Models

DGX agent

arXiv:2604.08395v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have greatly enhanced the integration of visual perception and linguistic reasoning, driving rapid prog

researcharxiv-cs-cv
10 Apr 2026
Safety

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

DGX agent

arXiv:2604.06628v1 Announce Type: new Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit t

safetyarxiv-cs-ai
10 Apr 2026
Research

$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models

DGX agent

arXiv:2604.06260v1 Announce Type: cross Abstract: Test-time scaling investigates whether a fixed diffusion language model (DLM) can generate better outputs when given more inference compute, without a

researcharxiv-cs-ai
10 Apr 2026
Research

The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models

DGX agent

arXiv:2604.06374v1 Announce Type: cross Abstract: Latent reasoning via continuous chain-of-thoughts (Latent CoT) has emerged as a promising alternative to discrete CoT reasoning. Operating in continuo

researcharxiv-cs-lg
10 Apr 2026
Applications

TREASURE: The Visa Payment Foundation Model for High-Volume Transaction Understanding

DGX agent

arXiv:2511.19693v3 Announce Type: replace-cross Abstract: Payment networks form the backbone of modern commerce, generating high volumes of transaction records from daily activities. Properly modeling

applicationsarxiv-cs-ai
10 Apr 2026
Model Releases

Variational Feature Compression for Model-Specific Representations

DGX agent

arXiv:2604.06644v1 Announce Type: cross Abstract: As deep learning inference is increasingly deployed in shared and cloud-based settings, a growing concern is input repurposing, in which data submitte

model-releasesarxiv-cs-lg
10 Apr 2026
Applications

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment

DGX agent

arXiv:2604.08212v1 Announce Type: new Abstract: General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring preci

applicationsarxiv-cs-cv
10 Apr 2026
Model Releases

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

DGX agent

arXiv:2604.07705v1 Announce Type: new Abstract: Aerial vision-and-language navigation (Aerial VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and autonom

model-releasesarxiv-cs-ro
10 Apr 2026
Safety

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

DGX agent

arXiv:2604.08546v1 Announce Type: new Abstract: Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a

safetyarxiv-cs-cv
10 Apr 2026
Model Releases

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

DGX agent

arXiv:2608.12307v1 Announce Type: cross Abstract: Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forc

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

Quantization Damage Is Multiplicative, Not Additive

DGX agent

arXiv:2608.06564v1 Announce Type: cross Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt. What nobody can say is which of the model's

model-releasesarxiv-cs-cl
10 Aug 2026
Model Releases

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools

DGX agent

arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Mo

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials

DGX agent

arXiv:2509.21079v2 Announce Type: replace Abstract: Foundation models have shown remarkable capabilities in various domains, but their performance on complex, multimodal engineering problems remains l

model-releasesarxiv-cs-cl
4 Aug 2026
← Previous
1…7475767778…1030
Next →