AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

86,428Total entries
1Added by human
86,427Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,016 results
29 May 2026

Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection

Model ReleasesDGX agent

arXiv:2605.29901v1 Announce Type: cross Abstract: Large language models (LLMs) can detect software vulnerabilities, but how do they actually identify vulnerable code? We address this question using me

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning

Model ReleasesDGX agent

arXiv:2605.29339v1 Announce Type: new Abstract: With the rapid advancement of multimodal large language models (MLLMs), models have demonstrated increasingly powerful multimodal capabilities. However,

Do Language Models Track Entities Across State Changes?

ResearchDGX agent

arXiv:2605.30233v1 Announce Type: cross Abstract: Entity tracking (ET), the ability to keep track of states, is a fundamental skill that underlies complex reasoning. An increasing amount of work inves

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR

Model ReleasesDGX agent

arXiv:2605.29637v1 Announce Type: new Abstract: Large language models recall knowledge reliably in English but often fail on the same query posed in a lower-resourced language -- a crosslingual consis

EvoMD-LLM: Learning the Language of Species Evolution in Reactive Molecular Dynamics

SafetyDGX agent

arXiv:2605.29394v1 Announce Type: new Abstract: While large language models (LLMs) excel at static scientific reasoning, they struggle to model the temporal structure of dynamic physical processes. We

From General Vision to Reliable Traversability Estimation: Adapting Vision Foundation Models for Unstructured Outdoor Environments

SafetyDGX agent

arXiv:2605.29565v1 Announce Type: new Abstract: Vision-based approaches have become the dominant paradigm for traversability estimation in unstructured outdoor environments, typically adapting vision

GeoMag: Geometric-Aware Video Motion Magnification via State Space Model

ApplicationsDGX agent

arXiv:2605.29762v1 Announce Type: new Abstract: Video Motion Magnification (VMM) reveals imperceptible dynamics but often suffers from structural inconsistencies under complex geometric transformation

Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning

SafetyDGX agent

arXiv:2605.29661v1 Announce Type: new Abstract: Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object

GPS-Enhanced Tourist Mobility Modeling with Seasonal Spatial Priors and LLM-Based Activity Chain Generation

ResearchDGX agent

arXiv:2605.29578v1 Announce Type: new Abstract: Tourist mobility poses a distinct challenge for urban transportation planning. Unlike resident commuting, tourist travel is largely non-routine, attract

Hallucination Detection-Guided Preference Optimization for Clinical Summarization

Model ReleasesDGX agent

arXiv:2605.28910v1 Announce Type: cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect stat

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.29782v1 Announce Type: cross Abstract: Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state va

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models

AgentsDGX agent

arXiv:2605.29625v1 Announce Type: new Abstract: The topic of Co-creation, i.e., AI agents interacting with humans to generate outputs (e.g., art), has gained significant attention recently. However, m

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

Model ReleasesDGX agent

arXiv:2605.29473v1 Announce Type: cross Abstract: Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond inf

Just dropped 🧑‍🍳 Fixed version of DeepSeek-V4-Pro-NVFP4 by @NVIDIAAI https://huggingface.co/nvidia/DeepSeek-V4-Pro-NVFP4

Model ReleasesDGX agent

NVIDIA has released a corrected version of DeepSeek-V4-Pro-NVFP4, a quantized model variant optimized for NVIDIA hardware using NV-FP4 (NVIDIA's 4-bit floating-point format). The model is available on

Leak@k: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding

Model ReleasesDGX agent

arXiv:2511.04934v3 Announce Type: replace Abstract: Unlearning in large language models (LLMs) is critical for regulatory compliance and for building ethical generative AI systems that avoid producing

Masked Diffusion Vision-Language Models for Temporal Action Localization

ResearchDGX agent

arXiv:2605.29858v1 Announce Type: new Abstract: Temporal action localization (TAL) requires recognizing the target event and localizing its start and end times precisely in untrimmed videos. Recent vi

MMTM: Tri-Modal Topic Modeling for Long-Form Video via Similarity-Gated Fusion

ResearchDGX agent

arXiv:2605.29765v1 Announce Type: new Abstract: We introduce MMTM, a modular pipeline for topic discovery in long-form video that integrates speech recognition, audio and visual embeddings, and BERTop

MVP-Shapley: Feature-based Modeling for Evaluating the Most Valuable Player in Basketball

ResearchDGX agent

arXiv:2506.04602v4 Announce Type: replace-cross Abstract: The burgeoning growth of the esports and multiplayer online gaming community has highlighted the critical importance of evaluating the Most Va

Permutation-Invariant Spectral Learning via Dyson Diffusion

SafetyDGX agent

arXiv:2510.08535v2 Announce Type: replace-cross Abstract: Diffusion models are central to generative modeling and have been adapted to graphs by diffusing adjacency matrix representations. The challen

Realistic honeypot evaluations for scheming propensity

Model ReleasesDGX agent

arXiv:2605.29729v1 Announce Type: new Abstract: We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming

ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for Zero-Shot Traffic Signal Control

SafetyDGX agent

arXiv:2605.29425v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promise in traffic signal control (TSC). However, its reliance on predefined states limits responsiveness to obser

SAHG: Sector-Anisotropic Hyperbolic Graph Model for Social Bot Detection

SafetyDGX agent

arXiv:2605.30166v1 Announce Type: cross Abstract: LLM-driven social bots can generate fluent, human-like text, reducing the discriminative advantage of content-based detection alone. However, coordina

SigmaMedStat: Temporal Signal Modeling for ICU False Alarm Reduction

SafetyDGX agent

arXiv:2605.29236v1 Announce Type: new Abstract: Alarm fatigue in intensive care units (ICUs) is a well documented patient safety crisis. Clinical monitors generate 350 or more alarms per patient per d

Small Agent Group is the Future of Digital Health

Model ReleasesDGX agent

arXiv:2602.08013v2 Announce Type: replace Abstract: The rapid adoption of large language models (LLMs) in digital health has been driven by a 'scaling-first' philosophy, i.e., the assumption that clin

The Trust Paradox: How CS Researchers Engage LLM Leaderboards

Model ReleasesDGX agent

arXiv:2605.28966v1 Announce Type: new Abstract: Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite kno

Towards Continuous-time Causal Foundation Models

ResearchDGX agent

arXiv:2605.28880v1 Announce Type: new Abstract: Extending discrete-time causal Prior-data Fitted Networks for time series to continuous time invites writing the mechanism as a stochastic differential

28 May 2026

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates

Model ReleasesDGX agent

arXiv:2605.28440v1 Announce Type: new Abstract: DPO has become a widely adopted alternative to RLHF for aligning LLMs with human preferences, eliminating the need for a separate reward model or RL loo

AI in SRE: Where and how Google is deploying agentic AI to improve operations

Model ReleasesDGX agent

Since its inception over 20 years ago, Google has used Site Reliability Engineering (SRE) to keep services like Search, Gmail, Maps, YouTube and Google Cloud reliable and highly available, adhering to

Applications of temporal graph learning for predicting the dynamics of biological systems

ResearchDGX agent

arXiv:2605.28659v1 Announce Type: new Abstract: Biological foundation models have shown strong performance in single-cell representation learning by applying transformer architectures directly to gene

Claude Opus 4.8: 'a modest but tangible improvement'

Model ReleasesDGX agent

Anthropic shipped Claude Opus 4.8 today. My favourite thing about it is this note in the release announcement: Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor. Ther

Debate Helps Weak Judges Reward Stronger Models

ResearchDGX agent

arXiv:2605.27483v1 Announce Type: cross Abstract: Despite theoretical promise, debate as a scalable oversight protocol has produced mixed empirical results: gains in some settings, and null effects in

DebFilter: Eradicating Biases Stashed in Value

SafetyDGX agent

arXiv:2605.28167v1 Announce Type: new Abstract: Text-to-image diffusion models, which are theoretically equivalent to score-based generative models, generate images through a multi-step denoising proc

FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation

Model ReleasesDGX agent

arXiv:2605.27849v1 Announce Type: cross Abstract: Despite rapid progress in LLM-based code generation, existing models are predominantly trained on imperative languages, leaving functional programming

Heterogeneous Multi-Agent Modeling for Measurement and Network Analysis of the Data Service Market

AgentsDGX agent

arXiv:2605.27433v1 Announce Type: cross Abstract: With the increasing complexity of collaboration among various social entities and user demands, the factors affecting the stable development of the da

Hybrid Neural World Models

ResearchDGX agent

arXiv:2605.28317v1 Announce Type: cross Abstract: Neural surrogates promise large speedups over classical solvers for physical dynamics but fail silently at sharp dynamical events such as shocks, fron

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

Model ReleasesDGX agent

arXiv:2605.27984v1 Announce Type: cross Abstract: Speech language models (SpeechLMs) have achieved substantial progress by extending large language models (LLMs) to the speech modality. However, Speec

Learning the Error Patterns of Language Models

TutorialsDGX agent

arXiv:2605.28328v1 Announce Type: cross Abstract: When generating outputs for domains with specific validity constraints (e.g., a program should compile), LLMs often fail in a small number of focused

Mathematical Modelling of Ethical AI Use in Higher Education: A Coordination Game Framework for Future-Facing Learning

SafetyDGX agent

arXiv:2605.27400v1 Announce Type: cross Abstract: The rapid uptake of generative artificial intelligence (AI) in higher education is reshaping assessment practices and intensifying concerns around aca

Measuring Massive Multitask Chinese Understanding

ApplicationsDGX agent

arXiv:2304.12986v3 Announce Type: replace-cross Abstract: The development of large-scale Chinese language models is flourishing, yet there is a lack of corresponding capability assessments. Therefore,

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation

Model ReleasesDGX agent

arXiv:2605.28035v1 Announce Type: new Abstract: In recent years, Multi-Talker Audio-Video Generation (MTAVG) models have shown promising performance on fundamental metrics such as lip-sync and audio-v

Multi-Adapter Representation Interventions via Energy Calibration

Model ReleasesDGX agent

arXiv:2605.28722v1 Announce Type: new Abstract: Representation intervention has emerged as a promising paradigm for aligning large language models toward desired behaviors without modifying model weig

MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation

Model ReleasesDGX agent

arXiv:2605.28579v1 Announce Type: new Abstract: Large language models (LLMs) have recently advanced text-driven 3D generation, yet Text-to-CAD remains far from supporting industrial product design. Ex

SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter

ApplicationsDGX agent

arXiv:2605.28084v1 Announce Type: cross Abstract: Laughter is a complex social signal that conveys communicative intent beyond amusement. While prior work has focused on isolated laughter analysis tas

The Alignment Floor: When Persona Customization Is Safe

Model ReleasesDGX agent

arXiv:2605.27382v1 Announce Type: cross Abstract: A key promise of pluralistic AI is behavioral adaptation: persona prompts like 'be creative' or 'be thorough' let systems respect diverse user values

TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

SafetyDGX agent

arXiv:2605.27690v1 Announce Type: new Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long be

27 May 2026

ADRD-Bench: A Preliminary LLM Benchmark for Alzheimer's Disease and Related Dementias

Model ReleasesDGX agent

arXiv:2602.11460v2 Announce Type: replace Abstract: Large language models (LLMs) have shown great potential for healthcare applications. However, existing evaluation benchmarks provide minimal coverag

Benchmark Leakage Trap: Can We Trust LLM-based Recommendation?

Model ReleasesDGX agent

arXiv:2602.13626v3 Announce Type: replace Abstract: The expanding integration of Large Language Models (LLMs) into recommender systems poses critical challenges to evaluation reliability. This paper i

Boosting Knowledge Graph Foundation Models via Enhanced Negative Sampling

ResearchDGX agent

arXiv:2605.27023v1 Announce Type: new Abstract: Knowledge graphs (KGs) have become the core backbone of numerous downstream tasks such as question answering and recommender systems. However, despite a

Conceptual Schema Inference for Tabular Datasets using Large Language Models

ApplicationsDGX agent

arXiv:2509.04632v2 Announce Type: replace-cross Abstract: Large collections of tabular data from data lakes, web tables and open data portals often originate from heterogeneous sources, leading to rep

Constraint acquisition needs better benchmarks

Model ReleasesDGX agent

arXiv:2605.26279v1 Announce Type: new Abstract: Constraint Acquisition (CA) and related research on the validation and enhancement of Mathematical Programming (MP) models from domain knowledge artifac

DunbaaBERT: From Sacrifice to Semantics

Model ReleasesDGX agent

arXiv:2605.26935v1 Announce Type: new Abstract: Large language models have achieved strong performance across many NLP tasks, yet Urdu remains comparatively underexplored due to limited resources and

EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation

SafetyDGX agent

arXiv:2605.26785v1 Announce Type: cross Abstract: Post-trained LLMs are often optimized to align responses with human preferences, making them safe, polite, and conversationally appropriate. In advers

FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentation

AgentsDGX agent

arXiv:2605.27178v1 Announce Type: cross Abstract: We address the challenging task of 3D object segmentation in complex scene point clouds without relying on any scene-level human annotations during tr

Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction

Model ReleasesDGX agent

arXiv:2605.26230v1 Announce Type: new Abstract: Multi-view 3D reconstruction has achieved remarkable progress with the advent of feed-forward 3D reconstruction models. However, these models are typica

I think Anthropic and OpenAI have found product-market fit

Model ReleasesDGX agent

Anthropic are strongly rumored to be about to have their first profitable quarter. Stories are circulating of companies surprised at how expensive their LLM bills are becoming from usage by their staf

MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning

ResearchDGX agent

arXiv:2605.26567v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode evidence-based decision logic that clinicians apply by evaluating patient variables, conditional criteria, an

MTL-FNO: A Lightweight Multi-Task Fourier Neural Operator for Sparse Field Reconstruction

Model ReleasesDGX agent

arXiv:2605.26718v1 Announce Type: new Abstract: Efficient onboard multi-field sparse reconstruction is essential for the autonomous operation of aerospace vehicles. While existing deep learning models

Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals

Model ReleasesDGX agent

arXiv:2605.26999v1 Announce Type: new Abstract: Prompt injection poses a critical threat to the safe deployment of large language models, yet existing detection approaches are typically evaluated unde

Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context Prefilling

Model ReleasesDGX agent

arXiv:2604.18103v2 Announce Type: replace Abstract: Prefilling computational costs pose a significant bottleneck for Large Language Models (LLMs) and Large Multimodal Models (LMMs) in long-context set

Tensormesh taps Nvidia, AMD and CoreWeave for funding to fix AI model memory problems

HardwareDGX agent

Tensormesh Inc. has hit upon a way to make artificial intelligence inference more efficient by eliminating the need for redundant computations, and its technology is so convincing that several of AI i

← Previous
1…258259260261262…1034
Next →