AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

86,510Total entries
1Added by human
86,509Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,082 results
29 May 2026

Temporal Stability and Few-Shot Prompting in Math Task Assessment

Model ReleasesDGX agent

arXiv:2605.30151v1 Announce Type: new Abstract: As AI tools become increasingly integrated into educational contexts, questions arise about both their stability over time and their responsiveness to p

The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF

Model ReleasesDGX agent

arXiv:2605.29491v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in agentic and retrieval-augmented generation (RAG) systems, where they must execute user-specifi

Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.29894v1 Announce Type: new Abstract: Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual task

When, why, and how do diffusion posterior samplers fail? A finite-sample lens

ResearchDGX agent

arXiv:2605.30330v1 Announce Type: new Abstract: Diffusion models have excellent capacity to model complex distributions of natural data, which has made them a popular and effective choice for posterio

28 May 2026

ATLAS: All-round Testing of Long-context Abilities across Scales

Model ReleasesDGX agent

arXiv:2605.28079v1 Announce Type: new Abstract: Long-context language models now advertise context windows up to millions of tokens, yet evaluations typically report a single length or a narrow task f

Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield Perspective

ResearchDGX agent

arXiv:2605.27476v1 Announce Type: cross Abstract: We characterize the pre-softmax attention matrix mathbf{QK^op} in transformers as an associative memory matrix encoding pairwise associations between

Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code

Model ReleasesDGX agent

arXiv:2603.24631v2 Announce Type: replace-cross Abstract: Code agents resolve 65-70% of SWE-bench Verified issues, but Pass@1 cannot tell us why the rest fail, and, as we show, capable-model failures

Dimensionality Reduction for Robust Federated Learning: A Theoretical Analysis and Convergence Guarantee

Model ReleasesDGX agent

arXiv:2605.28335v1 Announce Type: new Abstract: Federated Learning (FL) enables multiple clients to collaboratively train models without sharing raw data, but it is highly vulnerable to Byzantine atta

Do Agents Think Deeper? A Mechanistic Investigation of Layer-Wise Dynamics in Sequential Planning

Model ReleasesDGX agent

arXiv:2605.27935v1 Announce Type: new Abstract: Recent mechanistic studies suggest that large language models (LLMs) may utilize their depth inefficiently in standard single-turn tasks. Whether this s

Exploratory Experience Shapes the Geometry of Predictive Representations

Model ReleasesDGX agent

arXiv:2605.27929v1 Announce Type: cross Abstract: Active sensing links behavior and learning through an action-perception loop: actions determine the observations used to update internal predictive mo

Gradient Transformer: Learning to Generate Updates for LLMs

Model ReleasesDGX agent

arXiv:2605.27591v1 Announce Type: new Abstract: Many organizations lack computational resources to fine-tune large language models (LLMs) on private (unshareable) data for better utility, while fine-t

Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks

Model ReleasesDGX agent

arXiv:2605.27595v1 Announce Type: cross Abstract: Large Language Models (LLMs) are being rapidly adopted in agricultural imaging applications, ranging from crop interpretation to synthetic field image

ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment

SafetyDGX agent

arXiv:2605.27374v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) and diffusion models (DMs) have opened new possibilities for AI-generated content. Yet, pers

Intra-YOLO: A Small Object Detection Model for Caries and Molar-Incisor Hypomineralization in Intraoral Photography Based on Transfer Learning with Reinforcement Learning

ResearchDGX agent

arXiv:2605.28157v1 Announce Type: new Abstract: This study developed a computer-aided diagnosis (CAD) system for detecting caries and molar-incisor hypomineralization (MIH) in intraoral photographs. T

LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks

Model ReleasesDGX agent

arXiv:2605.27375v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context

LEIA: Learned Environment for Interactive Architected Materials

Model ReleasesDGX agent

arXiv:2605.28368v1 Announce Type: new Abstract: World models have enabled interactive exploration of game environments and robotic manipulation, but physical engineering remains beyond their reach: re

MIRA: A Bilingual Benchmark for Medical Information Response Audit

Model ReleasesDGX agent

arXiv:2605.28025v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide public-facing health information, yet existing safety evaluations overlook whether respons

Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors

ApplicationsDGX agent

arXiv:2605.27967v1 Announce Type: cross Abstract: Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), inclu

PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI

Model ReleasesDGX agent

arXiv:2605.27545v1 Announce Type: new Abstract: Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe text

Periodic RoPE for Infinite Context LLMs

Model ReleasesDGX agent

arXiv:2605.27980v1 Announce Type: cross Abstract: The ability to process ultra-long contexts is crucial for large language models (LLMs) to perform long-horizon tasks. While recent efforts have extend

Plan Before Search: Search Agents Need Plan

AgentsDGX agent

arXiv:2605.28354v1 Announce Type: new Abstract: Training large language models as retrieval-augmented reasoning agents typically combines reinforcement learning with an SFT cold start distilled from a

Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval

Model ReleasesDGX agent

arXiv:2605.11325v2 Announce Type: replace-cross Abstract: Every major benchmark for LLM memory systems, LoCoMo foremost, measures whether a model answered correctly, not whether the memory system retr

The Abstraction Gap in Vision-Language Causal Reasoning

Model ReleasesDGX agent

arXiv:2605.28779v1 Announce Type: new Abstract: Vision-language models (VLMs) generate fluent causal explanations, but current evaluations cannot distinguish linguistic plausibility from faithful caus

Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution

Model ReleasesDGX agent

arXiv:2605.28000v1 Announce Type: cross Abstract: Large language model agents are increasingly expected to perform operational work: calling APIs, manipulating files, assembling workflows, and acting

Understanding Generalization and Forgetting in In-Context Continual Learning

Model ReleasesDGX agent

arXiv:2605.28705v1 Announce Type: new Abstract: In-context learning (ICL) derives its power from enabling Large Language Models to adapt to new tasks via prompt-based reasoning alone, entirely bypassi

VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading

SafetyDGX agent

arXiv:2605.28818v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-la

Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation

Model ReleasesDGX agent

arXiv:2605.28187v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as scholar recommenders, shaping who is seen as an expert in academia. Existing audits remain Engli

27 May 2026

Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines

SafetyDGX agent

arXiv:2605.26442v1 Announce Type: cross Abstract: Much of the alignment tuning literature is organized around optimization objectives, while the construction of alignment data is often treated implici

Can LLMs Introspect? A Reality Check

ResearchDGX agent

arXiv:2605.26242v1 Announce Type: new Abstract: Can large language models detect and report their own internal states? A number of studies have argued that the answer to this question is yes. We argue

DEI: Diversity in Evolutionary Inference for Quality-Diversity Search

Model ReleasesDGX agent

arXiv:2605.27130v1 Announce Type: cross Abstract: We present DEI: Diversity in Evolutionary Inference, a distributed Quality-Diversity (QD) search framework that assigns heterogeneous large language m

Device Context Protocol: A Compact, Safety-First Architecture for LLM-Driven Control of Constrained Devices

Model ReleasesDGX agent

arXiv:2605.26159v1 Announce Type: cross Abstract: Large language models are increasingly used as orchestrators of external tools via the Model Context Protocol (MCP), but MCP is built for software ser

Diffuse to Detect: Generative Diffusion Models for Unsupervised IC Anomaly Detection

ResearchDGX agent

arXiv:2605.26468v1 Announce Type: cross Abstract: Latent defect screening is challenged by extremely low failure rates, high-dimensional test data, and absence of labeled anomalies. We propose the fir

EdgeFlow: Edge-Map Augmented VLM-Based Flowchart Processing for Industrial Requirements Engineering

Model ReleasesDGX agent

arXiv:2605.27332v1 Announce Type: cross Abstract: Flowcharts are widely used in industrial requirements, but usually remain embedded as static images. Vision Language Models (VLMs) show promise in the

Interpretability and Generalization Bounds for Learning Spatial Physics

Model ReleasesDGX agent

arXiv:2506.15199v3 Announce Type: replace Abstract: While there are many applications of ML to scientific problems that look promising, visuals can be deceiving. Using numerical analysis techniques, w

Learning to Adapt SFT Data for Better Reasoning Generalization

ResearchDGX agent

arXiv:2605.26924v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress, with post-training playing a crucial role in enhancing their reasoning capabilities. Amo

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

Model ReleasesDGX agent

arXiv:2605.26546v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents enable AI models to autonomously operate smartphones on behalf of users. However, most existing systems foc

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

Model ReleasesDGX agent

arXiv:2605.26485v1 Announce Type: cross Abstract: We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-vi

On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning

Model ReleasesDGX agent

arXiv:2605.27083v1 Announce Type: new Abstract: Counterfactual tuning (CFT) has emerged as a promising paradigm for Large Language Model (LLM) unlearning by training models to generate alternative fic

ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research

Model ReleasesDGX agent

arXiv:2601.21008v3 Announce Type: replace-cross Abstract: Operations Research practitioners debug infeasible models through an iterative process: inspecting Irreducible Infeasible Subsystems ( IIS), i

PILOT: A Data-Free Continual Learning Approach for Real-Time Semantic Segmentation via Boundary Guidance

TutorialsDGX agent

arXiv:2605.27128v1 Announce Type: new Abstract: Real-time semantic segmentation models offer an excellent balance between accuracy and inference speed. However, deploying these models in dynamic real

PRBench: A Standardized Probabilistic Robustness Benchmark

Model ReleasesDGX agent

arXiv:2511.01724v3 Announce Type: replace Abstract: Deep learning models are notoriously vulnerable to imperceptible perturbations. Most existing research centers on adversarial robustness (AR), which

Scheduled Style Injection: Expanding the Style-Content Pareto Frontier in Training-Free Diffusion-based Style Transfer

Model ReleasesDGX agent

arXiv:2605.26538v1 Announce Type: new Abstract: Style transfer with pre-trained diffusion models has advanced rapidly, but a core question remains underexplored: where in the model should style inject

The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?

Model ReleasesDGX agent

arXiv:2605.27176v1 Announce Type: new Abstract: Knowledge graphs (KGs) can provide structured scientific context to language models, but it remains unclear which graph facts actually shape the generat

The Daily Dose: Workflow-Integrated Large Language Model Automation for Clinical Summarization and Trial Identification in Radiation Oncology

ResearchDGX agent

arXiv:2605.26346v1 Announce Type: new Abstract: Objective: To describe the design and early clinical evaluation of The Daily Dose (TDD), an LLM-driven, automated clinical summarization and clinical-tr

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

Model ReleasesDGX agent

arXiv:2605.27141v1 Announce Type: new Abstract: Large language models (LLMs) have evolved into interactive agents that collaborate with users in real-world tasks. Effective collaboration in such setti

26 May 2026

A Comprehensive Dataset for Human vs. AI Generated Text Detection

Model ReleasesDGX agent

arXiv:2510.22874v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authentic

A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training

Model ReleasesDGX agent

arXiv:2605.24006v1 Announce Type: cross Abstract: Pipeline parallelism is a key technique for distributed training of large language models because it reduces per-device parameter and activation memor

Advancing Graph Few-Shot Learning via In-Context Learning

Model ReleasesDGX agent

arXiv:2605.24410v1 Announce Type: new Abstract: Graph few-shot learning, which aims to classify nodes from novel classes with only a few labeled examples, is a widely studied problem in graph learning

Asking LLMs to Verify First is Almost Free Lunch

Model ReleasesDGX agent

arXiv:2511.21734v2 Announce Type: replace-cross Abstract: To enhance the reasoning capabilities of Large Language Models (LLMs) without high costs of training, nor extensive test-time sampling, we int

CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation

Model ReleasesDGX agent

arXiv:2605.25378v1 Announce Type: cross Abstract: Customized image editing aims to equip pre-trained diffusion models with specific visual effects using limited paired data, typically via Low-Rank Ada

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

Model ReleasesDGX agent

arXiv:2605.24279v1 Announce Type: new Abstract: A frontier language model's acknowledged 'helpful programming assistant' persona does not survive long agentic-coding sessions in the deployment regime

D^2-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing

Model ReleasesDGX agent

arXiv:2605.25893v1 Announce Type: new Abstract: Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring

Distributionally Robust Transfer Learning with Structurally Missing Covariates, with Application to Cross-National Cardiac Arrest Prediction

Model ReleasesDGX agent

arXiv:2605.24212v1 Announce Type: cross Abstract: Deploying clinical prediction models across healthcare systems often fails when key training covariates are unavailable at deployment and labeled outc

Double Triangle Annotation: A Scalable Human-in-the-Loop Framework for High-Precision Historical Document Annotation

Model ReleasesDGX agent

arXiv:2605.25781v1 Announce Type: new Abstract: Evaluating structured-information extraction from historical documents at scale requires high-precision ground-truth annotations, yet traditional manual

FairJudge: Abstention-Aware Multimodal Judges for Fairness and Alignment Evaluation in Text-to-Image Models

SafetyDGX agent

arXiv:2510.22827v3 Announce Type: replace-cross Abstract: Evaluating text-to-image (T2I) systems requires judging not only whether an image matches a prompt, but also whether socially salient attribut

HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space

Model ReleasesDGX agent

arXiv:2509.22299v3 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) architectures in large language models (LLMs) deliver exceptional performance and reduced inference costs compared to

HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing

Model ReleasesDGX agent

arXiv:2605.24687v1 Announce Type: cross Abstract: Text-to-Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal bi

KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI

Model ReleasesDGX agent

arXiv:2510.02327v2 Announce Type: replace-cross Abstract: Real-time speech-to-speech (S2S) models excel at generating natural, low-latency conversational responses but often lack deep knowledge and se

Lattice theory and algebraic models for deep convolutional learning based on mathematical morphology

ResearchDGX agent

arXiv:2605.24608v1 Announce Type: new Abstract: We develop a rigorous algebraic framework for deep convolutional architectures, CNNs, ResNets, and encoder--decoder networks such as UNet, grounded in l

LLMs Show No Signs Of Individuated Metacognition

ResearchDGX agent

arXiv:2605.24299v1 Announce Type: new Abstract: Confidence-weighted routing, selective abstention, and ensemble weighting all assume that a model's stated confidence is informative about its capabilit

← Previous
1…281282283284285…1035
Next →