AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
48,975 results
Model Releases

Constructing Industrial-Scale Optimization Modeling Benchmark

DGX agent

arXiv:2602.10450v2 Announce Type: replace-cross Abstract: Optimization modeling underpins decision-making in logistics, manufacturing, energy, and finance, yet translating natural-language requirement

model-releasesarxiv-cs-ai
27 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning

DGX agent

arXiv:2605.26292v1 Announce Type: cross Abstract: Parameter-efficient adaptation of vision-language foundation models is crucial for precise multimodal understanding of biomedical images, yet existing

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Modeling Dynamic Mixtures of Time-Delay Systems from Streaming Time Series

DGX agent

arXiv:2605.26191v1 Announce Type: cross Abstract: This research addresses the problem of adaptive modeling in time-series data streams with clear input-output relationships. This problem is challengin

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding

DGX agent

arXiv:2605.26320v1 Announce Type: cross Abstract: The application of generalist multimodal models (GMMs) to specialized scientific domains remains limited due to the scarcity of comprehensive domain-s

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling

DGX agent

arXiv:2605.26322v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to infer others' knowledge, intentions, and emotions, is commonly evaluated in large language models (LLMs) using end-

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

A Large-Scale Dataset and Benchmark: Do Protein-Ligand Models Learn Binding Sites or Just Binding Likelihood?

DGX agent

arXiv:2605.24045v1 Announce Type: cross Abstract: Protein-ligand modeling underpins computational drug discovery and molecular design. Existing protein-ligand benchmarks typically evaluate whether a p

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Decision-Making with Lightweight Confidence-Aware Language Model for Autonomous Driving

DGX agent

arXiv:2605.25393v1 Announce Type: new Abstract: Large Language Models (LLMs) and Multimodal LLMs (MLLMs) have demonstrated immense potential in autonomous driving (AD) by offering human-like reasoning

model-releasesarxiv-cs-ro
26 May 2026
Model Releases

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

DGX agent

arXiv:2605.26038v1 Announce Type: cross Abstract: Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objec

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Emotional intelligence in large language models is fragmented across perception, cognition, and interaction

DGX agent

arXiv:2605.24686v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly integrated into emotionally sensitive domains, the structural integrity of their emotional intelligence

model-releasesarxiv-cs-ai
26 May 2026
Research

Large Language Model Selection with Limited Annotations

DGX agent

arXiv:2605.24981v1 Announce Type: new Abstract: Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations o

researcharxiv-cs-cl
26 May 2026
Research

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

DGX agent

arXiv:2509.07961v2 Announce Type: replace Abstract: We develop new experimental paradigms for measuring welfare in language models. We compare verbal reports of models about their preferences with pre

researcharxiv-cs-ai
26 May 2026
Model Releases

SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors

DGX agent

arXiv:2502.11167v5 Announce Type: replace-cross Abstract: Neural surrogate models are powerful and efficient tools in data mining. Meanwhile, large language models (LLMs) have demonstrated remarkable

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Toward a Benchmark for Controllable Simulation of Imperfect Students with Large Language Models

DGX agent

arXiv:2605.25601v1 Announce Type: cross Abstract: Teacher education requires deliberate practice with learners who exhibit identifiable strengths, weaknesses, and partial mastery. Large language model

model-releasesarxiv-cs-ai
26 May 2026
Safety

Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment

DGX agent

arXiv:2509.04445v2 Announce Type: replace Abstract: Recent AI trends seek to align AI models to learned human-centric objectives, such as personal preferences, utility, or societal values. Using stand

safetyarxiv-cs-lg
26 May 2026
Model Releases

When Mean CE Fails: Median CE Can Better Track Language Model Quality

DGX agent

arXiv:2605.24667v1 Announce Type: new Abstract: Mean cross-entropy is the standard validation metric for language models, but it can fail to track model quality during training. We examine this in two

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks

DGX agent

arXiv:2605.23243v1 Announce Type: cross Abstract: We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM

model-releasesarxiv-cs-ai
25 May 2026
Safety

Diffusion and Flow Matching Models for Tabular Data: A Survey

DGX agent

arXiv:2502.17119v2 Announce Type: replace-cross Abstract: Deep generative models have made rapid progress in image, text, audio, and video generation, and are increasingly being applied to structured

safetyarxiv-cs-ai
25 May 2026
Model Releases

DreamerNLplus: Interpretable Modeling of Mental Health Dynamics from Social Media Timelines using Hybrid Rule-Based and RAG Methods

DGX agent

arXiv:2605.23052v1 Announce Type: cross Abstract: We present DreamerNLplus, a hybrid framework for modeling mental health dynamics from social media timelines in the CLPsych 2026 shared task. Our syst

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Evaluating Large Language Models in a Complex Hidden Role Game

DGX agent

arXiv:2605.22826v1 Announce Type: cross Abstract: Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments.

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Operator Learning for Reconstructing Flow Fields from Sparse Measurements: a Language Model Approach

DGX agent

arXiv:2605.23712v1 Announce Type: cross Abstract: Reconstructing flow fields from sparse measurements is a fundamental problem in fluid mechanics with broad implications for modeling, control, and des

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

DGX agent

arXiv:2605.23019v1 Announce Type: new Abstract: Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other compon

model-releasesarxiv-cs-lg
25 May 2026
Local Ai

F-TIS: Harnessing Diverse Models in Collaborative GRPO

DGX agent

arXiv:2605.22537v1 Announce Type: new Abstract: Reinforcement learning methods such as GRPO have seen great popularity in LLM post-training. In GRPO, models produce completions to a set of prompts, wh

local-aiarxiv-cs-lg
23 May 2026
Model Releases

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

DGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

DGX agent

arXiv:2605.22080v1 Announce Type: new Abstract: We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF material

model-releasesarxiv-cs-cv
22 May 2026
Safety

Unifying Masked Diffusion Models with Various Generation Orders and Beyond

DGX agent

arXiv:2602.02112v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) are a potential alternative to autoregressive models (ARMs) for language generation, but generation quality dep

safetyarxiv-cs-cl
22 May 2026
Research

A Rigorous, Tractable Measure of Model Complexity

DGX agent

arXiv:2605.21167v1 Announce Type: cross Abstract: An accurate assessment of a model's complexity is crucial for topics such as interpretation, generalization, and model selection. However, most existi

researcharxiv-cs-lg
21 May 2026
Model Releases

CardioBench: Do Echocardiography Foundation Models Generalize Beyond the Lab?

DGX agent

arXiv:2510.00520v2 Announce Type: replace Abstract: Foundation models are reshaping medical imaging, yet their application in echocardiography remains limited, hindered by a heavy reliance on private

model-releasesarxiv-cs-cv
21 May 2026
Safety

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

DGX agent

arXiv:2602.02304v2 Announce Type: replace-cross Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning w

safetyarxiv-cs-lg
21 May 2026
Research

Corrected Integrated Laplace Approximation for Bayesian Inference in Latent Gaussian Models

DGX agent

arXiv:2605.20345v1 Announce Type: cross Abstract: Latent Gaussian models (LGMs) are a popular class of Bayesian hierarchical models that include Gaussian processes, as well as certain spatial models a

researcharxiv-cs-lg
21 May 2026
Research

Machine-Learning-Enhanced Non-Invasive Testing for MASLD Fibrosis: Shallow-Deep Neural Networks Versus FIB-4, Tabular Foundation Models, and Large Language Models

DGX agent

arXiv:2605.20523v1 Announce Type: new Abstract: Advanced fibrosis is a major determinant of liver-related morbidity in metabolic dysfunction-associated steatotic liver disease (MASLD). FIB-4 is widely

researcharxiv-cs-lg
21 May 2026
Model Releases

PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models

DGX agent

arXiv:2605.20873v1 Announce Type: cross Abstract: Planning is a fundamental capability for large language models (LLMs) because such complex tasks require models to coordinate goals, constraints, reso

model-releasesarxiv-cs-lg
21 May 2026
Model Releases

TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health

DGX agent

arXiv:2605.21295v1 Announce Type: new Abstract: Longitudinal passive sensing enables continuous health prediction, yet models often fail under cross-dataset distribution shifts. Traditional ML overfit

model-releasesarxiv-cs-lg
21 May 2026
Research

UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models

DGX agent

arXiv:2504.13109v2 Announce Type: replace Abstract: Flow matching models have emerged as a strong alternative to diffusion models, but existing inversion and editing methods designed for diffusion are

researcharxiv-cs-cv
21 May 2026
Safety

VLANeXt: Recipes for Building Strong VLA Models

DGX agent

arXiv:2602.18532v2 Announce Type: replace Abstract: Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding fro

safetyarxiv-cs-cv
21 May 2026
Model Releases

Disentangling generalization and memorization in large language models using chess

DGX agent

arXiv:2601.16823v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit remarkable capabilities, yet it remains unclear to what extent these reflect sophisticated recall or genu

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior

DGX agent

arXiv:2510.14261v2 Announce Type: replace Abstract: We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for interve

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks

DGX agent

arXiv:2605.19957v1 Announce Type: cross Abstract: World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single str

model-releasesarxiv-cs-ai
20 May 2026
Agents

A Mechanistic Model for Collective Motion from Sensorimotor Regularities

DGX agent

arXiv:2605.16522v1 Announce Type: new Abstract: Collective behavior in animals has long been modeled through self-propelled particle models, which reproduce striking group-level phenomena through abst

agentsarxiv-cs-ro
19 May 2026
Model Releases

DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models

DGX agent

arXiv:2601.11895v3 Announce Type: replace-cross Abstract: DevBench is a telemetry-driven benchmark designed to evaluate Large Language Models (LLMs) on realistic code completion tasks. It includes 1,8

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models

DGX agent

arXiv:2605.17070v1 Announce Type: new Abstract: While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on que

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

GIM: Evaluating models via tasks that integrate multiple cognitive domains

DGX agent

arXiv:2605.18663v1 Announce Type: new Abstract: As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA, HLE) or remo

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Machine Unlearning for Masked Diffusion Language Models

DGX agent

arXiv:2605.18253v1 Announce Type: cross Abstract: Recent masked diffusion language models (MDLMs), such as LLaDA and Dream, have achieved performance comparable to autoregressive large language models

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA

DGX agent

arXiv:2605.17932v1 Announce Type: cross Abstract: Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus primarily on autoregressive archite

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference

DGX agent

arXiv:2605.16360v1 Announce Type: cross Abstract: Efficient long-context inference in Large Language Models (LLMs) is severely constrained by the Key-Value (KV) cache memory wall, yet existing pruning

model-releasesarxiv-cs-ai
19 May 2026
Safety

Real-Time Aligned Reward Model beyond Semantics

DGX agent

arXiv:2601.22664v4 Announce Type: replace Abstract: Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique for aligning large language models (LLMs) with human preferences, yet it is

safetyarxiv-cs-ai
19 May 2026
Safety

Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations

DGX agent

arXiv:2605.16651v1 Announce Type: new Abstract: Explanation mechanisms are increasingly used to support transparency and trust in vision-language models (VLMs), particularly in settings where model de

safetyarxiv-cs-cv
19 May 2026
Research

SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models

DGX agent

arXiv:2605.17985v1 Announce Type: cross Abstract: We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. While model compression is essential

researcharxiv-cs-ai
19 May 2026
Model Releases

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models

DGX agent

arXiv:2605.18160v1 Announce Type: cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrati

model-releasesarxiv-cs-ai
19 May 2026
← Previous
1…3637383940…1021
Next →