AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,860 results
28 May 2026

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

Model ReleasesDGX agent

arXiv:2508.21046v3 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high com

Do Models Know Why They Changed Their Mind? Interpretability and Faithfulness of Chain-of-Thought Under Knowledge Conflict

Model ReleasesDGX agent

arXiv:2605.27773v1 Announce Type: cross Abstract: When a language model sees a document contradicting its training knowledge, it must choose: follow the document or trust itself. Prior work proved thi

JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2601.01627v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety

models being conscious would be harmful for humanity. it would encroach on our status and dignity. it would limit the type of things we can …

ApplicationsDGX agent

models being conscious would be harmful for humanity. it would encroach on our status and dignity. it would limit the type of things we can do with them and use them for. it would vastly accelerate hu

XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

Model ReleasesDGX agent

arXiv:2506.22726v4 Announce Type: replace Abstract: Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hin

27 May 2026

AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference

Model ReleasesDGX agent

arXiv:2512.11280v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved remarkable performance across a wide range of tasks, but their increasing parameter sizes significantly s

Alignment Makes Language Models Normative, Not Descriptive

SafetyDGX agent

arXiv:2603.17218v2 Announce Type: replace-cross Abstract: Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed

Beyond Questions: Evaluating What Large Language Models (Actually) Know

Model ReleasesDGX agent

arXiv:2605.26937v1 Announce Type: cross Abstract: Parametric knowledge in large language models (LLMs) is a cornerstone of their success, yet remains poorly understood. Existing knowledge benchmarks t

Constructing Industrial-Scale Optimization Modeling Benchmark

Model ReleasesDGX agent

arXiv:2602.10450v2 Announce Type: replace-cross Abstract: Optimization modeling underpins decision-making in logistics, manufacturing, energy, and finance, yet translating natural-language requirement

Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning

Model ReleasesDGX agent

arXiv:2605.26292v1 Announce Type: cross Abstract: Parameter-efficient adaptation of vision-language foundation models is crucial for precise multimodal understanding of biomedical images, yet existing

Modeling Dynamic Mixtures of Time-Delay Systems from Streaming Time Series

Model ReleasesDGX agent

arXiv:2605.26191v1 Announce Type: cross Abstract: This research addresses the problem of adaptive modeling in time-series data streams with clear input-output relationships. This problem is challengin

MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding

Model ReleasesDGX agent

arXiv:2605.26320v1 Announce Type: cross Abstract: The application of generalist multimodal models (GMMs) to specialized scientific domains remains limited due to the scarcity of comprehensive domain-s

OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling

Model ReleasesDGX agent

arXiv:2605.26322v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to infer others' knowledge, intentions, and emotions, is commonly evaluated in large language models (LLMs) using end-

26 May 2026

A Large-Scale Dataset and Benchmark: Do Protein-Ligand Models Learn Binding Sites or Just Binding Likelihood?

Model ReleasesDGX agent

arXiv:2605.24045v1 Announce Type: cross Abstract: Protein-ligand modeling underpins computational drug discovery and molecular design. Existing protein-ligand benchmarks typically evaluate whether a p

Decision-Making with Lightweight Confidence-Aware Language Model for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.25393v1 Announce Type: new Abstract: Large Language Models (LLMs) and Multimodal LLMs (MLLMs) have demonstrated immense potential in autonomous driving (AD) by offering human-like reasoning

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

Model ReleasesDGX agent

arXiv:2605.26038v1 Announce Type: cross Abstract: Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objec

Emotional intelligence in large language models is fragmented across perception, cognition, and interaction

Model ReleasesDGX agent

arXiv:2605.24686v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly integrated into emotionally sensitive domains, the structural integrity of their emotional intelligence

Large Language Model Selection with Limited Annotations

ResearchDGX agent

arXiv:2605.24981v1 Announce Type: new Abstract: Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations o

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

ResearchDGX agent

arXiv:2509.07961v2 Announce Type: replace Abstract: We develop new experimental paradigms for measuring welfare in language models. We compare verbal reports of models about their preferences with pre

Qwen3.7 Max now available in Go - text only - 1M context - smartest model in the Qwen family to date

Model ReleasesDGX agent

Alibaba's Qwen team has released Qwen3.7 Max, their most advanced model to date, now available for use with Go programming language support. The model features a 1 million token context window and is

SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors

Model ReleasesDGX agent

arXiv:2502.11167v5 Announce Type: replace-cross Abstract: Neural surrogate models are powerful and efficient tools in data mining. Meanwhile, large language models (LLMs) have demonstrated remarkable

Toward a Benchmark for Controllable Simulation of Imperfect Students with Large Language Models

Model ReleasesDGX agent

arXiv:2605.25601v1 Announce Type: cross Abstract: Teacher education requires deliberate practice with learners who exhibit identifiable strengths, weaknesses, and partial mastery. Large language model

Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment

SafetyDGX agent

arXiv:2509.04445v2 Announce Type: replace Abstract: Recent AI trends seek to align AI models to learned human-centric objectives, such as personal preferences, utility, or societal values. Using stand

When Mean CE Fails: Median CE Can Better Track Language Model Quality

Model ReleasesDGX agent

arXiv:2605.24667v1 Announce Type: new Abstract: Mean cross-entropy is the standard validation metric for language models, but it can fail to track model quality during training. We examine this in two

25 May 2026

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks

Model ReleasesDGX agent

arXiv:2605.23243v1 Announce Type: cross Abstract: We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM

Diffusion and Flow Matching Models for Tabular Data: A Survey

SafetyDGX agent

arXiv:2502.17119v2 Announce Type: replace-cross Abstract: Deep generative models have made rapid progress in image, text, audio, and video generation, and are increasingly being applied to structured

DreamerNLplus: Interpretable Modeling of Mental Health Dynamics from Social Media Timelines using Hybrid Rule-Based and RAG Methods

Model ReleasesDGX agent

arXiv:2605.23052v1 Announce Type: cross Abstract: We present DreamerNLplus, a hybrid framework for modeling mental health dynamics from social media timelines in the CLPsych 2026 shared task. Our syst

Evaluating Large Language Models in a Complex Hidden Role Game

Model ReleasesDGX agent

arXiv:2605.22826v1 Announce Type: cross Abstract: Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments.

Operator Learning for Reconstructing Flow Fields from Sparse Measurements: a Language Model Approach

Model ReleasesDGX agent

arXiv:2605.23712v1 Announce Type: cross Abstract: Reconstructing flow fields from sparse measurements is a fundamental problem in fluid mechanics with broad implications for modeling, control, and des

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

Model ReleasesDGX agent

arXiv:2605.23019v1 Announce Type: new Abstract: Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other compon

23 May 2026

F-TIS: Harnessing Diverse Models in Collaborative GRPO

Local AiDGX agent

arXiv:2605.22537v1 Announce Type: new Abstract: Reinforcement learning methods such as GRPO have seen great popularity in LLM post-training. In GRPO, models produce completions to a set of prompts, wh

22 May 2026

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

Model ReleasesDGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

Model ReleasesDGX agent

arXiv:2605.22080v1 Announce Type: new Abstract: We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF material

Unifying Masked Diffusion Models with Various Generation Orders and Beyond

SafetyDGX agent

arXiv:2602.02112v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) are a potential alternative to autoregressive models (ARMs) for language generation, but generation quality dep

21 May 2026

A Rigorous, Tractable Measure of Model Complexity

ResearchDGX agent

arXiv:2605.21167v1 Announce Type: cross Abstract: An accurate assessment of a model's complexity is crucial for topics such as interpretation, generalization, and model selection. However, most existi

CardioBench: Do Echocardiography Foundation Models Generalize Beyond the Lab?

Model ReleasesDGX agent

arXiv:2510.00520v2 Announce Type: replace Abstract: Foundation models are reshaping medical imaging, yet their application in echocardiography remains limited, hindered by a heavy reliance on private

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

SafetyDGX agent

arXiv:2602.02304v2 Announce Type: replace-cross Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning w

Corrected Integrated Laplace Approximation for Bayesian Inference in Latent Gaussian Models

ResearchDGX agent

arXiv:2605.20345v1 Announce Type: cross Abstract: Latent Gaussian models (LGMs) are a popular class of Bayesian hierarchical models that include Gaussian processes, as well as certain spatial models a

Machine-Learning-Enhanced Non-Invasive Testing for MASLD Fibrosis: Shallow-Deep Neural Networks Versus FIB-4, Tabular Foundation Models, and Large Language Models

ResearchDGX agent

arXiv:2605.20523v1 Announce Type: new Abstract: Advanced fibrosis is a major determinant of liver-related morbidity in metabolic dysfunction-associated steatotic liver disease (MASLD). FIB-4 is widely

PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models

Model ReleasesDGX agent

arXiv:2605.20873v1 Announce Type: cross Abstract: Planning is a fundamental capability for large language models (LLMs) because such complex tasks require models to coordinate goals, constraints, reso

TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health

Model ReleasesDGX agent

arXiv:2605.21295v1 Announce Type: new Abstract: Longitudinal passive sensing enables continuous health prediction, yet models often fail under cross-dataset distribution shifts. Traditional ML overfit

UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models

ResearchDGX agent

arXiv:2504.13109v2 Announce Type: replace Abstract: Flow matching models have emerged as a strong alternative to diffusion models, but existing inversion and editing methods designed for diffusion are

VLANeXt: Recipes for Building Strong VLA Models

SafetyDGX agent

arXiv:2602.18532v2 Announce Type: replace Abstract: Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding fro

20 May 2026

Disentangling generalization and memorization in large language models using chess

Model ReleasesDGX agent

arXiv:2601.16823v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit remarkable capabilities, yet it remains unclear to what extent these reflect sophisticated recall or genu

Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior

Model ReleasesDGX agent

arXiv:2510.14261v2 Announce Type: replace Abstract: We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for interve

What we learned testing 7 models under the same agent harness

Model ReleasesDGX agent

Model swaps look like configuration changes, but they behave more like product migrations. A new model may be cheaper, faster, easier to get capacity for, or stronger on public benchmarks.... The post

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks

Model ReleasesDGX agent

arXiv:2605.19957v1 Announce Type: cross Abstract: World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single str

19 May 2026

A Mechanistic Model for Collective Motion from Sensorimotor Regularities

AgentsDGX agent

arXiv:2605.16522v1 Announce Type: new Abstract: Collective behavior in animals has long been modeled through self-propelled particle models, which reproduce striking group-level phenomena through abst

DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models

Model ReleasesDGX agent

arXiv:2601.11895v3 Announce Type: replace-cross Abstract: DevBench is a telemetry-driven benchmark designed to evaluate Large Language Models (LLMs) on realistic code completion tasks. It includes 1,8

EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.17070v1 Announce Type: new Abstract: While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on que

GIM: Evaluating models via tasks that integrate multiple cognitive domains

Model ReleasesDGX agent

arXiv:2605.18663v1 Announce Type: new Abstract: As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA, HLE) or remo

I remember when people were saying 'It's useless to open-source big models because nobody will be able to run them fast'....

Model ReleasesDGX agent

I remember when people were saying 'It's useless to open-source big models because nobody will be able to run them fast'.... Cerebras is now running Kimi K2.6 – a trillion parameter model – in enterpr

Machine Unlearning for Masked Diffusion Language Models

Model ReleasesDGX agent

arXiv:2605.18253v1 Announce Type: cross Abstract: Recent masked diffusion language models (MDLMs), such as LLaDA and Dream, have achieved performance comparable to autoregressive large language models

Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA

Model ReleasesDGX agent

arXiv:2605.17932v1 Announce Type: cross Abstract: Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus primarily on autoregressive archite

ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference

Model ReleasesDGX agent

arXiv:2605.16360v1 Announce Type: cross Abstract: Efficient long-context inference in Large Language Models (LLMs) is severely constrained by the Key-Value (KV) cache memory wall, yet existing pruning

Real-Time Aligned Reward Model beyond Semantics

SafetyDGX agent

arXiv:2601.22664v4 Announce Type: replace Abstract: Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique for aligning large language models (LLMs) with human preferences, yet it is

Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations

SafetyDGX agent

arXiv:2605.16651v1 Announce Type: new Abstract: Explanation mechanisms are increasingly used to support transparency and trust in vision-language models (VLMs), particularly in settings where model de

SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models

ResearchDGX agent

arXiv:2605.17985v1 Announce Type: cross Abstract: We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. While model compression is essential

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.18160v1 Announce Type: cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrati

18 May 2026

Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging

Model ReleasesDGX agent

arXiv:2505.21698v3 Announce Type: replace Abstract: Vision-language foundation models achieve promising performance in natural image classification, yet their direct application to medical imaging is

← Previous
1…3940414243…998
Next →