AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,612 results
2 Jun 2026

XAI-SOH-FL: Enhancing SOH-FL with Adaptive Aggregation and Explainable AI for Intrusion Detection in Heterogeneous IoT

Model ReleasesDGX agent

arXiv:2606.00134v1 Announce Type: cross Abstract: Intrusion Detection Systems (IDS) in Internet of Things (IoT) environments face significant challenges due to data heterogeneity, lack of labeled data

1 Jun 2026

Auditing LLM Benchmarks with Item Response Theory

Model ReleasesDGX agent

arXiv:2605.30504v1 Announce Type: new Abstract: LLM benchmark labels are frozen at release and silently propagated into downstream benchmarks, errors and all. We introduce an Item Response Theory-base

Bandwidth Allocation with Device Partitioning for Federated Learning over Industrial IoT networks

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2605.30892v1 Announce Type: new Abstract: We consider a federated learning (FL) system in which Industrial Internet-of-Things (IIoT) devices collaboratively train a global model over wireless ch

Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

Model ReleasesDGX agent

arXiv:2605.31086v1 Announce Type: new Abstract: In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the under

BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs

Model ReleasesDGX agent

arXiv:2605.30900v1 Announce Type: new Abstract: Current multimodal models handle static image recognition well, but intuitive physical reasoning remains a weakness. Predicting how objects will move an

Convergence of Steepest Descent and Adam under Non-Uniform Smoothness

Model ReleasesDGX agent

arXiv:2605.30648v1 Announce Type: new Abstract: Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

SafetyDGX agent

arXiv:2605.30590v1 Announce Type: cross Abstract: Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

Model ReleasesDGX agent

arXiv:2602.10809v2 Announce Type: replace Abstract: Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This

Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages

Model ReleasesDGX agent

arXiv:2605.30529v1 Announce Type: cross Abstract: Sentence-embedding models for semantic search are overwhelmingly developed and evaluated on English corpora. When applied to clinical retrieval in oth

Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation

Model ReleasesDGX agent

arXiv:2605.30984v1 Announce Type: cross Abstract: Modern 3D medical vision-language models (VLMs) can generate fluent radiology-style text while exhibit critically low pathology detection and output d

Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs

TutorialsDGX agent

arXiv:2510.00419v2 Announce Type: replace Abstract: Zeroth-order optimizers have recently emerged as an attractive approach for fine-tuning large language models (LLMs), as they avoid backpropagation

MIMO: Multilingual Information Retrieval via Monolingual Objectives

Model ReleasesDGX agent

arXiv:2605.31171v1 Announce Type: cross Abstract: Multilingual Information Retrieval (MLIR) reflects real-world search environments in which queries and relevant documents may appear in different lang

NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMs

Model ReleasesDGX agent

arXiv:2505.17595v4 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve impressive performance across domains but face significant challenges when deployed on consumer-grade GPU

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference

Model ReleasesDGX agent

arXiv:2510.07651v2 Announce Type: replace-cross Abstract: Large language models (LLMs) with extended context windows enable powerful applications but impose significant memory overhead, as caching all

REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge

SafetyDGX agent

arXiv:2603.17145v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as automated evaluators that assign numeric scores to model outputs, a paradigm known a

Rethinking Sparse Mixture of Experts from a Unified Perspective

ResearchDGX agent

arXiv:2503.22996v3 Announce Type: replace Abstract: Sparse Mixture of Experts (SMoE) models scale the capacity of models while maintaining constant computational overhead. SMoE methods fall into two c

SAW-Bench: Learning Situated Awareness in the Real World

Model ReleasesDGX agent

arXiv:2602.16682v2 Announce Type: replace Abstract: A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over pos

Scaling Multi-Hop Training Data via Graph-Constrained Path Selection

Model ReleasesDGX agent

arXiv:2605.31238v1 Announce Type: new Abstract: Endowing large language models with compositional reasoning over specialized documents requires multi-hop training data at scale, where such data rarely

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?

TutorialsDGX agent

arXiv:2605.30557v1 Announce Type: cross Abstract: Spatial reasoning is a fundamental capability for vision-language models (VLMs) deployed in real-world environments. However, visual observations are

SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching

SafetyDGX agent

arXiv:2605.30729v1 Announce Type: new Abstract: Schema matching is a fundamental step in integrating heterogeneous data sources. While Pre-trained Language Models (PLMs) have revolutionized this task

Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail

TutorialsDGX agent

arXiv:2605.31244v1 Announce Type: new Abstract: Neural scaling laws describe predictable power-law relationships between model size, dataset size, compute, and performance. While these laws guide the

The Surface You Test Is Not the Surface That Breaks

Model ReleasesDGX agent

arXiv:2605.30454v1 Announce Type: cross Abstract: Tool-augmented LLM agents are vulnerable to prompt injection: a third party who controls part of the agent's context can plant instructions that the a

UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities

Model ReleasesDGX agent

arXiv:2603.23160v2 Announce Type: replace Abstract: Benchmarking large language models (LLMs) and agents in multi-turn interactive scenarios is essential for understanding their practical capabilities

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

Model ReleasesDGX agent

arXiv:2509.24901v4 Announce Type: replace-cross Abstract: Although probing frozen models has become a standard evaluation paradigm, self-supervised learning in audio defaults to fine-tuning when pursu

29 May 2026

Active Learning for Machine Learning Driven Molecular Dynamics

Model ReleasesDGX agent

arXiv:2509.17208v3 Announce Type: replace Abstract: Machine-learned coarse-grained (CG) potentials are fast, but degrade over time when simulations reach under-sampled bio-molecular conformations, and

AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.29643v1 Announce Type: new Abstract: Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence d

Agree! was talking about this with @havoyan just a few days ago. That's also the reason why so much of the value has been accruing to the fr…

IndustryDGX agent

Agree! was talking about this with @havoyan just a few days ago. That's also the reason why so much of the value has been accruing to the frontier models in my opinion (cc @GavinSBaker) because if you

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

Model ReleasesDGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

Model ReleasesDGX agent

arXiv:2605.30188v1 Announce Type: cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibr

Connecting Independently Trained Modes via Layer-Wise Connectivity

Model ReleasesDGX agent

arXiv:2505.02604v5 Announce Type: replace Abstract: Empirical studies have shown that continuous low-loss paths can be constructed between independently trained neural network models. This phenomenon,

Deep Adaptive Dimension Reduction for Bayesian Inference in Inverse Problems

Model ReleasesDGX agent

arXiv:2605.29373v1 Announce Type: new Abstract: Solving high-dimensional PDE-governed inverse problems is often challenging due to complex non-Gaussian posterior distributions, expensive forward model

Deep Psychovisual Image Representations

TutorialsDGX agent

arXiv:2605.29260v1 Announce Type: new Abstract: Psychovisual models suggest human vision decouples low-level feature extraction from higher cognition by first forming intermediate abstractions. In con

DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

Model ReleasesDGX agent

arXiv:2605.29256v1 Announce Type: cross Abstract: Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality

ESPO: Early-Stopping Proximal Policy Optimization

Model ReleasesDGX agent

arXiv:2605.29860v1 Announce Type: cross Abstract: When a large language model under reinforcement learning commits a wrong reasoning step early in a trajectory, standard algorithms force it to keep ge

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

ResearchDGX agent

arXiv:2601.08654v2 Announce Type: replace-cross Abstract: Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with huma

GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation

SafetyDGX agent

arXiv:2605.28995v1 Announce Type: new Abstract: Recent approaches integrating vision-language models (VLMs) as prompt encoders for generative model conditioning typically rely on expensive end-to-end

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

Model ReleasesDGX agent

arXiv:2605.29107v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a gr

Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence

TutorialsDGX agent

arXiv:2605.30093v1 Announce Type: new Abstract: Foundation features from self-supervised vision models and text-to-image diffusion models have proven effective for semantic correspondence estimation.

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

Model ReleasesDGX agent

arXiv:2605.29668v1 Announce Type: new Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the

How Braintrust turns customer requests into code with Codex

Model ReleasesDGX agent

Braintrust leverages OpenAI's Codex model to automatically convert customer requests and natural language specifications into functional code, streamlining the software development process. This appli

How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning

Model ReleasesDGX agent

arXiv:2602.02103v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning has become a central mechanism for eliciting multi-step reasoning in Large Language Models (LLMs). Yet recent

LiveSVG: Zero-Shot SVG Animation via Video Generation

Model ReleasesDGX agent

arXiv:2605.30174v1 Announce Type: new Abstract: We introduce LiveSVG, a zero-shot approach for generating Scalable Vector Graphics (SVG) animations using video diffusion models. Current SVG animation

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

Model ReleasesDGX agent

arXiv:2605.29737v1 Announce Type: cross Abstract: LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code t

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

Model ReleasesDGX agent

arXiv:2605.29738v1 Announce Type: cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

Model ReleasesDGX agent

arXiv:2605.29300v1 Announce Type: cross Abstract: Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses ar

NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs

Model ReleasesDGX agent

arXiv:2605.29716v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising non-autoregressive generative paradigm. Given the prohibitive computational cost of

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

Model ReleasesDGX agent

arXiv:2605.29685v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

Model ReleasesDGX agent

arXiv:2605.29676v1 Announce Type: new Abstract: Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default languag

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

Model ReleasesDGX agent

arXiv:2605.29833v1 Announce Type: new Abstract: As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdi

On the Construction and Implications of Low-Loss Valleys in LoRA-based Bayesian Inference

Model ReleasesDGX agent

arXiv:2605.29580v1 Announce Type: new Abstract: While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic

Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies

Model ReleasesDGX agent

arXiv:2605.30148v1 Announce Type: cross Abstract: Evolution Strategies (ES) has recently emerged as a competitive alternative to reinforcement learning (RL) for large language model (LLM) fine-tuning,

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

SafetyDGX agent

arXiv:2605.29582v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide pro

Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning

Model ReleasesDGX agent

arXiv:2605.29028v1 Announce Type: cross Abstract: Conditioned Sequence Models (CSMs) learn policies by treating return-to-go (RTG) as a control signal. However, existing CSMs often treat the RTGs as s

Robust and Efficient Guardrails with Latent Reasoning

Model ReleasesDGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

Model ReleasesDGX agent

arXiv:2605.29468v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (

Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge

Model ReleasesDGX agent

arXiv:2605.29402v1 Announce Type: cross Abstract: Understanding long-form egocentric videos remains challenging for multimodal large language models (MLLMs) due to limited context length and insuffici

Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

SafetyDGX agent

arXiv:2605.29032v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably

Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.30344v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have achieved impressive performance across many tasks, yet prior studies report unsatisfactory perform

Took me a while to figure out what all the ESMFold2 rage was about. At first, the benchmarking data didn't look super remarkable to me but i…

Model ReleasesDGX agent

Took me a while to figure out what all the ESMFold2 rage was about. At first, the benchmarking data didn't look super remarkable to me but it turns there are many impressive aspects: - Fully open sour

TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation

Model ReleasesDGX agent

arXiv:2605.29656v1 Announce Type: new Abstract: Evaluating open-ended outputs from large language models (LLMs) remains challenging due to the absence of ground truth. Existing metrics rely on final-a

← Previous
1…347348349350351…1044
Next →