AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,612 results
20 May 2026

A Reproducibility Analysis of PO4ISR: Diagnosing and Mitigating Semantic Drift in LLM-Based Session Recommendation

Model ReleasesDGX agent

arXiv:2605.18780v1 Announce Type: cross Abstract: Reasoning-based Large Language Models (LLMs) like PO4ISR have set new benchmarks in session-based recommendation. However, the reproducibility of thei

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls

Model ReleasesDGX agent

arXiv:2605.19728v1 Announce Type: new Abstract: Foundation video models produce visually impressive results, but their use in embodied AI remains limited because they are primarily trained on natural

AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.19260v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have recently emerged as promising backbones for GUI-agent models, where high-resolution GUI screenshots are introduced t

Block-Based Double Decoders

ResearchDGX agent

arXiv:2605.18807v1 Announce Type: cross Abstract: Encoder-decoder models offer substantial inference-time savings over decoder-only models, but their pretraining objectives suffer from sparse supervis

brain dump of how/why we use Evals to measure agents before & after shipping to prod 1. Good Evals simulate what our real users will do and …

Model ReleasesDGX agent

brain dump of how/why we use Evals to measure agents before & after shipping to prod 1. Good Evals simulate what our real users will do and encounter. They’re not really random benchmark tasks, they r

Cross-Paradigm Knowledge Distillation: A Comprehensive Study of Bidirectional Transfer Between Random Forests and Deep Neural Networks for Big Data Applications

ResearchDGX agent

arXiv:2605.19299v1 Announce Type: new Abstract: The exponential growth of big data has intensified the need for efficient and interpretable machine learning models that can handle diverse data charact

Diffusion and Flow-based Copulas: Forgetting and Remembering Dependencies

TutorialsDGX agent

arXiv:2509.19707v2 Announce Type: replace-cross Abstract: Copulas are a fundamental tool for modelling multivariate dependencies in data, forming the method of choice in diverse fields and application

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data

Model ReleasesDGX agent

arXiv:2605.19130v1 Announce Type: cross Abstract: Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal

ESLD (External Surrogate Latent Defense): A Latent-Space Architecture for Faster, Stronger Prompt-Injection Defense

SafetyDGX agent

arXiv:2605.18918v1 Announce Type: cross Abstract: Modern AI assistants are agentic. To answer a single user request, the underlying language model pulls in information from many sources, such as web s

FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data

Model ReleasesDGX agent

arXiv:2605.18936v1 Announce Type: cross Abstract: Social media text data are often used to train Machine Learning (ML) models to identify users exhibiting high-risk mental health behaviors. However, s

From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning

Model ReleasesDGX agent

arXiv:2605.19824v1 Announce Type: new Abstract: Recent attempts to support high-level scene interpretation and planning in Autonomous Vehicles (AVs) using ensembles of Large Language Models (LLMs) and

GRAB: A Risk Taxonomy--Grounded Benchmark for Unsupervised Topic Discovery in Financial Disclosures

Model ReleasesDGX agent

arXiv:2509.21698v2 Announce Type: replace Abstract: Risk categorization in 10-K risk disclosures matters for oversight and investment, yet no public benchmark evaluates unsupervised topic models for t

Hallucination as Exploit: Evidence-Carrying Multimodal Agents

Model ReleasesDGX agent

arXiv:2605.19192v1 Announce Type: new Abstract: Multimodal agents use screenshots, documents, and webpages to choose tool calls. When a false visual claim triggers a click, email, extraction, or trans

How Does Overparameterization Affect Machine Unlearning of Deep Neural Networks?

SafetyDGX agent

arXiv:2503.08633v2 Announce Type: replace Abstract: Machine unlearning is the task of updating a trained model to forget specific training data without retraining from scratch. In this paper, we inves

INSIGHTS: Demonstration-Based Summaries of Time Series Predictors

Local AiDGX agent

arXiv:2605.18849v1 Announce Type: cross Abstract: Explainability methods have progressed rapidly, but global explanations for time-series models remain underdeveloped, with most approaches focusing on

Learn-by-Wire Training Control Governance: Bounded Autonomous Training Under Stress for Stability and Efficiency

Model ReleasesDGX agent

arXiv:2605.19008v1 Announce Type: new Abstract: Modern language-model training is increasingly exposed to instability, degraded runs, and wasted compute, especially under aggressive learning-rate, sca

Measuring Safety Alignment Effects in Autonomous Security Agents

Model ReleasesDGX agent

arXiv:2605.19722v1 Announce Type: cross Abstract: Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Sin

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

Model ReleasesDGX agent

arXiv:2605.20183v1 Announce Type: new Abstract: Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However,

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

Model ReleasesDGX agent

arXiv:2605.18825v1 Announce Type: new Abstract: Prefix caching is a key optimization in Large Language Model (LLM) serving, reusing attention Key-Value (KV) states across requests with shared prompt p

Perceptual misalignment of texture representations in convolutional neural networks

Local AiDGX agent

arXiv:2604.01341v2 Announce Type: replace Abstract: Mathematical modeling of visual textures traces back to Julesz's intuition that texture perception in humans is based on local correlations between

✨ Personal AI is the next computing platform. AI is shifting from something you access to something you build with, locally, at the edge, an…

Model ReleasesDGX agent

✨ Personal AI is the next computing platform. AI is shifting from something you access to something you build with, locally, at the edge, and across systems. We’re unlocking new possibilities for deve

Precision Tracked Transformer via Kalman Filtering, Kriging and Process Noise

Model ReleasesDGX agent

arXiv:2605.18832v1 Announce Type: cross Abstract: The Transformer is the foundational building block of modern AI, yet offers no principled handling of uncertainty, which is prevalent in real applicat

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

Model ReleasesDGX agent

arXiv:2605.19382v1 Announce Type: new Abstract: Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluat

Reporting from Google I/O 2026 with the four biggest themes from one of the biggest AI labs in the world. 🎤 Voice AI as an interface Google…

Model ReleasesDGX agent

Reporting from Google I/O 2026 with the four biggest themes from one of the biggest AI labs in the world. 🎤 Voice AI as an interface Google and Samsung announced new Gemini-powered glasses with Gentle

Robustness and Regularization in Hierarchical Re-Basin

ResearchDGX agent

arXiv:2510.09174v3 Announce Type: replace Abstract: This paper takes a closer look at Git Re-Basin, an interesting new approach to merge trained models. We propose a hierarchical model merging scheme

STAR-PolyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision

Model ReleasesDGX agent

arXiv:2605.19338v1 Announce Type: cross Abstract: Frontier AI models and multi-agent systems have led to significant improvements in mathematical reasoning. However, for problems requiring extended, l

Streamlined Constraint Reasoning via CNN Pattern Recognition on Enumerated Solutions

Model ReleasesDGX agent

arXiv:2605.19895v1 Announce Type: new Abstract: Constraint programming practitioners accelerate hard problems through a layered set of techniques applied in order of risk. Standard hardening (symmetry

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding

Model ReleasesDGX agent

arXiv:2605.19866v1 Announce Type: new Abstract: Vision-Language Models (VLMs) parse documents end-to-end but frequently break down on layouts unlike those seen in training. We attribute this to a two-

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload

HardwareDGX agent

arXiv:2605.20179v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive (AR) models, offering better hardware utilization an

Toto 2.0: Time Series Forecasting Enters the Scaling Era

Model ReleasesDGX agent

arXiv:2605.20119v1 Announce Type: cross Abstract: We show that time series foundation models scale: a single training recipe produces reliable forecast-quality improvements from 4M to 2.5B parameters.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs

Model ReleasesDGX agent

arXiv:2605.19528v1 Announce Type: new Abstract: 3D localization in Multimodal Large Language Models (MLLMs), including 3D object detection and 3D visual grounding, is fundamentally limited by camera i

Towards Distillation Guarantees under Algorithmic Alignment for Combinatorial Optimization

SafetyDGX agent

arXiv:2605.20074v1 Announce Type: new Abstract: Distillation transfers knowledge from a large model trained on broad data to a smaller, more efficient model suitable for deployment. In structured pred

Trust or Abstain? A Self-Aware RAG Approach

Model ReleasesDGX agent

arXiv:2605.18792v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves large language models (LLMs) by incorporating external evidence, but it also introduces knowledge confli

ViroGym: Realistic Large-Scale Benchmarks for Evaluating Viral Proteins

Model ReleasesDGX agent

arXiv:2603.06740v2 Announce Type: replace-cross Abstract: Protein language models (pLMs) have shown strong potential for zero-shot prediction of missense variant effects, yet systematic benchmarking o

WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions

Model ReleasesDGX agent

arXiv:2510.09872v2 Announce Type: replace-cross Abstract: Training web agents to navigate complex, real-world websites requires them to master extit{subtasks} - short-horizon interactions on multiple

What Makes a Representation Good for Single-Cell Perturbation Prediction?

Model ReleasesDGX agent

arXiv:2605.19343v1 Announce Type: new Abstract: Single-cell perturbation modeling is fundamental for understanding and predicting cellular responses to genetic perturbations. However, existing approac

Worldlier: native support for 48 world languages and improved efficiency in non-European languages.

Model ReleasesDGX agent

Worldlier is a Cohere language model that natively supports 48 world languages with improved efficiency, particularly for non-European languages. This represents an expansion of language coverage beyo

Your Neighbors Know: Leveraging Local Neighborhoods for Backdoor Detection in Decentralized Learning

Local AiDGX agent

arXiv:2605.19969v1 Announce Type: new Abstract: Decentralized learning (DL) is an emerging machine learning paradigm where nodes collaboratively train models without a central server. However, the col

19 May 2026

A Machine with Short-Term, Episodic, and Semantic Memory Systems

Model ReleasesDGX agent

arXiv:2212.02098v5 Announce Type: replace Abstract: Inspired by the cognitive science theory of the explicit human memory systems, we have modeled an agent with short-term, episodic, and semantic memo

A More Word-like Image Tokenization for MLLMs

ResearchDGX agent

arXiv:2605.17954v1 Announce Type: cross Abstract: Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a seque

Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring

Model ReleasesDGX agent

arXiv:2605.16386v1 Announce Type: new Abstract: Multimodal large language models (LLMs) are increasingly explored as automated evaluators in clinical settings, yet their scoring behavior on ordinal cl

BESplit: Bias-Compensated Split Federated Learning with Evidential Aggregation

Model ReleasesDGX agent

arXiv:2605.17508v1 Announce Type: cross Abstract: Split Federated Learning (SFL) enables privacy-preserving collaborative training by partitioning models between clients and a server. However, under n

Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers

Model ReleasesDGX agent

arXiv:2605.16949v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers (DiTs) demonstrate that aligning noisy latent states with well-trained semantic features-as pioneered by Repre

Boundedly Rational Meta-Learning in Sequential Consumer Choice

Model ReleasesDGX agent

arXiv:2605.16532v1 Announce Type: new Abstract: Many consumer decisions are repeated choices under uncertainty. Standard models capture these decisions using Bayesian learning and dynamic programming:

CADS: Conformal Adaptive Decision System for Cost-Efficient Image Classification

ResearchDGX agent

arXiv:2605.16401v1 Announce Type: new Abstract: While high-capacity AI models have advanced state-of-the-art performance, their practical deployment is often hindered by high inference costs, environm

CAM-Bench: A Benchmark for Computational and Applied Mathematics in Lean

Model ReleasesDGX agent

arXiv:2605.17255v1 Announce Type: new Abstract: Formal theorem-proving benchmarks enable mechanically verifiable evaluation of mathematical reasoning in large language models. However, existing benchm

Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks

ResearchDGX agent

arXiv:2510.01782v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) should refuse to answer questions beyond their knowledge. This capability, which we term knowledge-aware refusal,

CasualSynth: Generating Structurally Sound Synthetic Data

Model ReleasesDGX agent

arXiv:2605.17528v1 Announce Type: cross Abstract: Large Language Models (LLMs) generate realistic synthetic data but offer no guarantee that their outputs respect the causal mechanisms governing the t

ClaHF: A Human Feedback-inspired Reinforcement Learning Framework for Improving Classification Tasks

SafetyDGX agent

arXiv:2605.17458v1 Announce Type: new Abstract: Text classification models are typically trained via supervised fine-tuning (SFT). However, SFT essentially performs behavior cloning from instance-wise

Closing the Gap at CRAC 2026: Two-Stage Adaptation for LLM-Based Multilingual Coreference Resolution

Model ReleasesDGX agent

arXiv:2605.16984v1 Announce Type: new Abstract: We present our submission to the LLM track of the 2026 Computational Models of Reference, Anaphora and Coreference (CRAC 2026) shared task. With an aver

ContractBench: Can LLM Agents Preserve Observation Contracts?

Model ReleasesDGX agent

arXiv:2605.17281v1 Announce Type: cross Abstract: Tool-augmented LLM agents call APIs whose intermediate outputs, such as presigned URLs, session tokens, and OAuth state parameters, are observation co

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.02979v2 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong potential in complex reasoning, yet their progress remains fundamentally constrained by reliance

DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs

Model ReleasesDGX agent

arXiv:2605.18498v1 Announce Type: cross Abstract: Expert specialization in Mixture-of-Experts (MoE) models remains poorly understood, with traditional evaluations conflating architectural load-balanci

Disappointing pricing trend with Gemini 3.5 Flash. 22.5x pricier than 2.0 Flash which came out 15 months ago (9.00 vs 0.40). Are Flash mod…

Model ReleasesDGX agent

Disappointing pricing trend with Gemini 3.5 Flash. 22.5x pricier than 2.0 Flash which came out 15 months ago (9.00 vs 0.40). Are Flash models supposed to get this much more expensive, or is Pro just b

Distributed Perceptron under Bounded Staleness, Partial Participation, and Noisy Communication

Model ReleasesDGX agent

arXiv:2601.10705v3 Announce Type: replace Abstract: We study a semi-asynchronous client-server perceptron trained via iterative parameter mixing (IPM-style averaging): clients run local perceptron upd

DriveSafer: End-to-End Autonomous Driving with Safety Guidance

Model ReleasesDGX agent

arXiv:2605.16737v1 Announce Type: cross Abstract: End-to-End (E2E) autonomous driving models have shown growing capability in recent years, with performance improving on increasingly challenging bench

DynMuon: A Dynamic Spectral Shaping View of Muon

Model ReleasesDGX agent

arXiv:2605.17109v1 Announce Type: cross Abstract: In recent years, Muon has emerged as the dominant method for training large language models, and transformers more broadly. The essential difference,

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks

ResearchDGX agent

arXiv:2603.11689v2 Announce Type: replace Abstract: Frontier Multimodal Large Language Models (MLLMs) exhibit remarkable capabilities in Visual-Language Comprehension (VLC) tasks. However, they are of

Federated Learning by Utility-Constrained Stochastic Aggregation for Improving Rational Participation

Local AiDGX agent

arXiv:2605.18020v1 Announce Type: new Abstract: Federated Learning (FL) algorithms implicitly assume that clients passively comply with server-side orchestration by sharing local model updates upon se

FedSDR: Federated Self-Distillation with Rectification

Local AiDGX agent

arXiv:2605.18028v1 Announce Type: cross Abstract: Federated fine-tuning of Large Language Models faces severe statistical heterogeneity. However, existing model-level defenses often overlook the root

← Previous
1…351352353354355…1044
Next →