AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,903 results
Model Releases

Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought

DGX agent

arXiv:2605.27764v1 Announce Type: cross Abstract: Recent segmentation models couple large language models (LLMs) with mask decoders to ground complex language expressions into masks, yet their instruc

model-releasesarxiv-cs-ai
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Claude’s new model is more ‘honest’ when it messes up

DGX agent

Anthropic is releasing Claude Opus 4.8 on Thursday, and the company is touting the model's 'honesty.' According to Anthropic, it trains 'all [its] models to be honest - for instance, to avoid making c

model-releasesthe-verge-ai
28 May 2026
Model Releases

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

DGX agent

arXiv:2508.21046v3 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high com

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Do Models Know Why They Changed Their Mind? Interpretability and Faithfulness of Chain-of-Thought Under Knowledge Conflict

DGX agent

arXiv:2605.27773v1 Announce Type: cross Abstract: When a language model sees a document contradicting its training knowledge, it must choose: follow the document or trust itself. Prior work proved thi

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models

DGX agent

arXiv:2601.01627v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety

model-releasesarxiv-cs-ai
28 May 2026
Applications

models being conscious would be harmful for humanity. it would encroach on our status and dignity. it would limit the type of things we can …

DGX agent

models being conscious would be harmful for humanity. it would encroach on our status and dignity. it would limit the type of things we can do with them and use them for. it would vastly accelerate hu

applicationsethan-mollick--x
28 May 2026
Model Releases

XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

DGX agent

arXiv:2506.22726v4 Announce Type: replace Abstract: Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hin

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference

DGX agent

arXiv:2512.11280v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved remarkable performance across a wide range of tasks, but their increasing parameter sizes significantly s

model-releasesarxiv-cs-cl
27 May 2026
Safety

Alignment Makes Language Models Normative, Not Descriptive

DGX agent

arXiv:2603.17218v2 Announce Type: replace-cross Abstract: Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed

safetyarxiv-cs-ai
27 May 2026
Model Releases

Beyond Questions: Evaluating What Large Language Models (Actually) Know

DGX agent

arXiv:2605.26937v1 Announce Type: cross Abstract: Parametric knowledge in large language models (LLMs) is a cornerstone of their success, yet remains poorly understood. Existing knowledge benchmarks t

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Constructing Industrial-Scale Optimization Modeling Benchmark

DGX agent

arXiv:2602.10450v2 Announce Type: replace-cross Abstract: Optimization modeling underpins decision-making in logistics, manufacturing, energy, and finance, yet translating natural-language requirement

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning

DGX agent

arXiv:2605.26292v1 Announce Type: cross Abstract: Parameter-efficient adaptation of vision-language foundation models is crucial for precise multimodal understanding of biomedical images, yet existing

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Modeling Dynamic Mixtures of Time-Delay Systems from Streaming Time Series

DGX agent

arXiv:2605.26191v1 Announce Type: cross Abstract: This research addresses the problem of adaptive modeling in time-series data streams with clear input-output relationships. This problem is challengin

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding

DGX agent

arXiv:2605.26320v1 Announce Type: cross Abstract: The application of generalist multimodal models (GMMs) to specialized scientific domains remains limited due to the scarcity of comprehensive domain-s

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling

DGX agent

arXiv:2605.26322v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to infer others' knowledge, intentions, and emotions, is commonly evaluated in large language models (LLMs) using end-

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

A Large-Scale Dataset and Benchmark: Do Protein-Ligand Models Learn Binding Sites or Just Binding Likelihood?

DGX agent

arXiv:2605.24045v1 Announce Type: cross Abstract: Protein-ligand modeling underpins computational drug discovery and molecular design. Existing protein-ligand benchmarks typically evaluate whether a p

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Decision-Making with Lightweight Confidence-Aware Language Model for Autonomous Driving

DGX agent

arXiv:2605.25393v1 Announce Type: new Abstract: Large Language Models (LLMs) and Multimodal LLMs (MLLMs) have demonstrated immense potential in autonomous driving (AD) by offering human-like reasoning

model-releasesarxiv-cs-ro
26 May 2026
Model Releases

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

DGX agent

arXiv:2605.26038v1 Announce Type: cross Abstract: Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objec

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Emotional intelligence in large language models is fragmented across perception, cognition, and interaction

DGX agent

arXiv:2605.24686v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly integrated into emotionally sensitive domains, the structural integrity of their emotional intelligence

model-releasesarxiv-cs-ai
26 May 2026
Research

Large Language Model Selection with Limited Annotations

DGX agent

arXiv:2605.24981v1 Announce Type: new Abstract: Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations o

researcharxiv-cs-cl
26 May 2026
Research

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

DGX agent

arXiv:2509.07961v2 Announce Type: replace Abstract: We develop new experimental paradigms for measuring welfare in language models. We compare verbal reports of models about their preferences with pre

researcharxiv-cs-ai
26 May 2026
Model Releases

Qwen3.7 Max now available in Go - text only - 1M context - smartest model in the Qwen family to date

DGX agent

Alibaba's Qwen team has released Qwen3.7 Max, their most advanced model to date, now available for use with Go programming language support. The model features a 1 million token context window and is

model-releasesqwen--x
26 May 2026
Model Releases

SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors

DGX agent

arXiv:2502.11167v5 Announce Type: replace-cross Abstract: Neural surrogate models are powerful and efficient tools in data mining. Meanwhile, large language models (LLMs) have demonstrated remarkable

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Toward a Benchmark for Controllable Simulation of Imperfect Students with Large Language Models

DGX agent

arXiv:2605.25601v1 Announce Type: cross Abstract: Teacher education requires deliberate practice with learners who exhibit identifiable strengths, weaknesses, and partial mastery. Large language model

model-releasesarxiv-cs-ai
26 May 2026
Safety

Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment

DGX agent

arXiv:2509.04445v2 Announce Type: replace Abstract: Recent AI trends seek to align AI models to learned human-centric objectives, such as personal preferences, utility, or societal values. Using stand

safetyarxiv-cs-lg
26 May 2026
Model Releases

When Mean CE Fails: Median CE Can Better Track Language Model Quality

DGX agent

arXiv:2605.24667v1 Announce Type: new Abstract: Mean cross-entropy is the standard validation metric for language models, but it can fail to track model quality during training. We examine this in two

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks

DGX agent

arXiv:2605.23243v1 Announce Type: cross Abstract: We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM

model-releasesarxiv-cs-ai
25 May 2026
Safety

Diffusion and Flow Matching Models for Tabular Data: A Survey

DGX agent

arXiv:2502.17119v2 Announce Type: replace-cross Abstract: Deep generative models have made rapid progress in image, text, audio, and video generation, and are increasingly being applied to structured

safetyarxiv-cs-ai
25 May 2026
Model Releases

DreamerNLplus: Interpretable Modeling of Mental Health Dynamics from Social Media Timelines using Hybrid Rule-Based and RAG Methods

DGX agent

arXiv:2605.23052v1 Announce Type: cross Abstract: We present DreamerNLplus, a hybrid framework for modeling mental health dynamics from social media timelines in the CLPsych 2026 shared task. Our syst

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Evaluating Large Language Models in a Complex Hidden Role Game

DGX agent

arXiv:2605.22826v1 Announce Type: cross Abstract: Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments.

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Operator Learning for Reconstructing Flow Fields from Sparse Measurements: a Language Model Approach

DGX agent

arXiv:2605.23712v1 Announce Type: cross Abstract: Reconstructing flow fields from sparse measurements is a fundamental problem in fluid mechanics with broad implications for modeling, control, and des

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

DGX agent

arXiv:2605.23019v1 Announce Type: new Abstract: Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other compon

model-releasesarxiv-cs-lg
25 May 2026
Local Ai

F-TIS: Harnessing Diverse Models in Collaborative GRPO

DGX agent

arXiv:2605.22537v1 Announce Type: new Abstract: Reinforcement learning methods such as GRPO have seen great popularity in LLM post-training. In GRPO, models produce completions to a set of prompts, wh

local-aiarxiv-cs-lg
23 May 2026
Model Releases

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

DGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

DGX agent

arXiv:2605.22080v1 Announce Type: new Abstract: We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF material

model-releasesarxiv-cs-cv
22 May 2026
Safety

Unifying Masked Diffusion Models with Various Generation Orders and Beyond

DGX agent

arXiv:2602.02112v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) are a potential alternative to autoregressive models (ARMs) for language generation, but generation quality dep

safetyarxiv-cs-cl
22 May 2026
Research

A Rigorous, Tractable Measure of Model Complexity

DGX agent

arXiv:2605.21167v1 Announce Type: cross Abstract: An accurate assessment of a model's complexity is crucial for topics such as interpretation, generalization, and model selection. However, most existi

researcharxiv-cs-lg
21 May 2026
Model Releases

CardioBench: Do Echocardiography Foundation Models Generalize Beyond the Lab?

DGX agent

arXiv:2510.00520v2 Announce Type: replace Abstract: Foundation models are reshaping medical imaging, yet their application in echocardiography remains limited, hindered by a heavy reliance on private

model-releasesarxiv-cs-cv
21 May 2026
Safety

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

DGX agent

arXiv:2602.02304v2 Announce Type: replace-cross Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning w

safetyarxiv-cs-lg
21 May 2026
Research

Corrected Integrated Laplace Approximation for Bayesian Inference in Latent Gaussian Models

DGX agent

arXiv:2605.20345v1 Announce Type: cross Abstract: Latent Gaussian models (LGMs) are a popular class of Bayesian hierarchical models that include Gaussian processes, as well as certain spatial models a

researcharxiv-cs-lg
21 May 2026
Research

Machine-Learning-Enhanced Non-Invasive Testing for MASLD Fibrosis: Shallow-Deep Neural Networks Versus FIB-4, Tabular Foundation Models, and Large Language Models

DGX agent

arXiv:2605.20523v1 Announce Type: new Abstract: Advanced fibrosis is a major determinant of liver-related morbidity in metabolic dysfunction-associated steatotic liver disease (MASLD). FIB-4 is widely

researcharxiv-cs-lg
21 May 2026
Model Releases

PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models

DGX agent

arXiv:2605.20873v1 Announce Type: cross Abstract: Planning is a fundamental capability for large language models (LLMs) because such complex tasks require models to coordinate goals, constraints, reso

model-releasesarxiv-cs-lg
21 May 2026
Model Releases

TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health

DGX agent

arXiv:2605.21295v1 Announce Type: new Abstract: Longitudinal passive sensing enables continuous health prediction, yet models often fail under cross-dataset distribution shifts. Traditional ML overfit

model-releasesarxiv-cs-lg
21 May 2026
Research

UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models

DGX agent

arXiv:2504.13109v2 Announce Type: replace Abstract: Flow matching models have emerged as a strong alternative to diffusion models, but existing inversion and editing methods designed for diffusion are

researcharxiv-cs-cv
21 May 2026
Safety

VLANeXt: Recipes for Building Strong VLA Models

DGX agent

arXiv:2602.18532v2 Announce Type: replace Abstract: Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding fro

safetyarxiv-cs-cv
21 May 2026
Model Releases

Disentangling generalization and memorization in large language models using chess

DGX agent

arXiv:2601.16823v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit remarkable capabilities, yet it remains unclear to what extent these reflect sophisticated recall or genu

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior

DGX agent

arXiv:2510.14261v2 Announce Type: replace Abstract: We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for interve

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

What we learned testing 7 models under the same agent harness

DGX agent

Model swaps look like configuration changes, but they behave more like product migrations. A new model may be cheaper, faster, easier to get capacity for, or stronger on public benchmarks.... The post

model-releasesarize-ai
20 May 2026
← Previous
1…4950515253…1248
Next →