AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,566 results
13 May 2026

A Switching System Theory of Q-Learning with Linear Function Approximation

Model ReleasesDGX agent

arXiv:2605.11021v1 Announce Type: new Abstract: This paper develops a switching-system interpretation of Q-learning with linear function approximation (LFA) based on the joint spectral radius (JSR). W

A Theoretical Analysis of Why Masked Diffusion Models Mitigate the Reversal Curse

Model ReleasesDGX agent

arXiv:2602.02133v2 Announce Type: replace-cross Abstract: Autoregressive language models (ARMs) suffer from the reversal curse: after learning ''A is B,'' they often fail on the reverse query ''B is A

ABRA: Agent Benchmark for Radiology Applications

Model ReleasesDGX agent

arXiv:2605.11224v1 Announce Type: new Abstract: Existing medical-agent benchmarks deliver imaging as pre-selected samples, never as an environment the agent must navigate. We introduce ABRA, a radiolo


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

Model ReleasesDGX agent

arXiv:2605.11398v1 Announce Type: cross Abstract: We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations.

Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference

Model ReleasesDGX agent

arXiv:2605.11581v1 Announce Type: new Abstract: When large language models (LLMs) serve real-time inference in commercial online advertising systems, end-to-end latency must be strictly bounded to the

Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness Predictions

Model ReleasesDGX agent

arXiv:2211.03524v2 Announce Type: replace Abstract: Modern Review Helpfulness Prediction systems are dependent upon multiple modalities, typically texts and images. Unfortunately, those contemporary a

Agent-Based Post-Hoc Correction of Agricultural Yield Forecasts

Model ReleasesDGX agent

arXiv:2605.12375v1 Announce Type: new Abstract: Accurate crop yield forecasting in commercial soft fruit production is constrained by the data available in typical commercial farm records, which lack

Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty

Model ReleasesDGX agent

arXiv:2605.11436v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed on long-horizon tasks in partially observable environments, where they must act while inferring a

AgentDisCo: Towards Disentanglement and Collaboration in Open-ended Deep Research Agents

Model ReleasesDGX agent

arXiv:2605.11732v1 Announce Type: cross Abstract: In this paper, we present AgentDisCo, a novel Disentangled and Collaborative agentic architecture that formulates deep research as an adversarial opti

Allegory of the Cave: Measurement-Grounded Vision-Language Learning

Model ReleasesDGX agent

arXiv:2605.11727v1 Announce Type: cross Abstract: Vision-language models typically reason over post-ISP RGB images, although RGB rendering can clip, suppress, or quantize sensor evidence before infere

An Empirical Study of Automating Agent Evaluation

Model ReleasesDGX agent

arXiv:2605.11378v1 Announce Type: new Abstract: Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive

Anthropic launches Claude for Small Business, featuring a host of automated services like bookkeeping functions, business insights, and tools for ad campaigns (Lucas Ropek/TechCrunch)

Model ReleasesDGX agent

Lucas Ropek / TechCrunch: Anthropic launches Claude for Small Business, featuring a host of automated services like bookkeeping functions, business insights, and tools for ad campaigns — Anthropic is

Anthropic unveils Claude Agent SDK credits for paid plans, which users can allocate for programmatic use of third-party agents like OpenClaw, starting June 15 (Carl Franzen/VentureBeat)

Model ReleasesDGX agent

Carl Franzen / VentureBeat: Anthropic unveils Claude Agent SDK credits for paid plans, which users can allocate for programmatic use of third-party agents like OpenClaw, starting June 15 — Good news,

Approximation Theory of Laplacian-Based Neural Operators for Reaction-Diffusion System

Model ReleasesDGX agent

arXiv:2605.12025v1 Announce Type: new Abstract: Neural operators provide a framework for learning solution operators of partial differential equations (PDEs), enabling efficient surrogate modeling for

As President Trump meets President Xi this week, a call to the American AI community: If your startup, lab, non-profit or company benefits f…

Model ReleasesDGX agent

As President Trump meets President Xi this week, a call to the American AI community: If your startup, lab, non-profit or company benefits from open international AI - especially Chinese (Deepseek, Qw

ASD-Bench: A Four-Axis Comprehensive Benchmark of AI Models for Autism Spectrum Disorder

Model ReleasesDGX agent

arXiv:2605.11091v1 Announce Type: new Abstract: Automated ASD screening tools remain limited by single-architecture evaluations, axis-restricted assessment, and near-exclusive focus on adult cohorts,

AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor

Model ReleasesDGX agent

arXiv:2601.05752v3 Announce Type: replace Abstract: We introduce AutoMonitor-Bench, the first benchmark designed to systematically evaluate the reliability of LLM-based misbehavior monitors across div

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding

Model ReleasesDGX agent

arXiv:2508.05269v2 Announce Type: replace Abstract: Understanding dynamic outdoor environments requires capturing complex object interactions and their evolution over time. LiDAR-based 4D point clouds

Backbone-Equated Diffusion OOD via Sparse Internal Snapshots

Model ReleasesDGX agent

arXiv:2605.11014v1 Announce Type: new Abstract: Fair comparison between diffusion-based OOD detectors is challenging, as conclusions can vary with backbone choice, corruption parameterization, and tes

BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

Model ReleasesDGX agent

arXiv:2605.12074v1 Announce Type: new Abstract: Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a sce

BEExformer: A Fast Inferencing Binarized Transformer with Early Exits

Model ReleasesDGX agent

arXiv:2412.05225v3 Announce Type: replace Abstract: Large Language Models (LLMs) based on transformers achieve cutting-edge results on a variety of applications. However, their enormous size and proce

Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training

Model ReleasesDGX agent

arXiv:2605.12483v1 Announce Type: new Abstract: In settings where labeled verifiable training data is the binding constraint, each checked example should be allocated carefully. The standard practice

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

Model ReleasesDGX agent

arXiv:2605.12413v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study

Beyond Parameter Aggregation: Semantic Consensus for Federated Fine-Tuning of LLMs

Model ReleasesDGX agent

arXiv:2605.11857v1 Announce Type: new Abstract: Federated fine-tuning of large language models is commonly formulated as a parameter aggregation problem. However, even parameter-efficient methods requ

Beyond Prediction: Interval Neural Networks for Uncertainty-Aware System Identification

Model ReleasesDGX agent

arXiv:2605.11460v1 Announce Type: new Abstract: System identification (SysID) is critical for modeling dynamical systems from experimental data, yet traditional approaches often fail to capture nonlin

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm

Model ReleasesDGX agent

arXiv:2605.12271v1 Announce Type: new Abstract: Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generat

Bin Latent Transformer (BiLT): A shift-invariant autoencoder for calibration-free spectral unmixing of turbid media

Model ReleasesDGX agent

arXiv:2605.11829v1 Announce Type: cross Abstract: The accurate recovery of constituent-level optical properties from integrating sphere measurements is a central analytical challenge in pharmaceutical

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models

Model ReleasesDGX agent

arXiv:2605.11726v1 Announce Type: new Abstract: Recently, reinforcement learning (RL) has been widely applied during post-training for diffusion large language models (dLLMs) to enhance reasoning with

Blumira launches Kindling pilot, an agentic SIEM investigation engine that cuts alert volume up to 50x

Model ReleasesDGX agent

Security operations platform startup Blumira Inc. today launched the pilot of Kindling, an agentic security information and event management investigation engine that the company says can reduce alert

我們開源了這顆星球🌎上速度最快的低成本 bm25 引擎。

Model ReleasesDGX agent

我們開源了這顆星球🌎上速度最快的低成本 bm25 引擎。 so we built psql_bm25s. exact BM25 retrieval. native Postgres access method. ~23x faster than pg_search on the standard benchmark. retrieval stops being a budget item. the

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

Model ReleasesDGX agent

arXiv:2605.12034v1 Announce Type: cross Abstract: Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evid

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents

Model ReleasesDGX agent

arXiv:2508.07642v3 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3

Breaking extit{Winner-Takes-All}: Cooperative Policy Optimization Improves Diverse LLM Reasoning

Model ReleasesDGX agent

arXiv:2605.11461v1 Announce Type: cross Abstract: Reinforcement learning with verifiers (RLVR) has become a central paradigm for improving LLM reasoning, yet popular group-based optimization algorithm

Built a local coding harness powered by Gemma 4. It runs locally, connects to my model backend, starts coding sessions, streams responses, a…

Model ReleasesDGX agent

Built a local coding harness powered by Gemma 4. It runs locally, connects to my model backend, starts coding sessions, streams responses, and uses tools through a CLI-style workflow. Still early, but

CAD-feature enhanced machine learning for manufacturing effort estimation on sheet metal bending parts

Model ReleasesDGX agent

arXiv:2605.12266v1 Announce Type: new Abstract: Graph-based machine learning has emerged as a promising approach for manufacturability analysis by learning directly from CAD models represented as Boun

Calibrated Multimodal Representation Learning with Missing Modalities

Model ReleasesDGX agent

arXiv:2511.12034v2 Announce Type: replace Abstract: Multimodal representation learning harmonizes distinct modalities by aligning them into a unified latent space. Recent research generalizes traditio

Caraman at SemEval-2026 Task 8: Three-Stage Multi-Turn Retrieval with Query Rewriting, Hybrid Search, and Cross-Encoder Reranking

Model ReleasesDGX agent

arXiv:2605.12028v1 Announce Type: new Abstract: We describe our system for SemEval-2026 Task 8 (MTRAGEval), participating in Task A (Retrieval) across four English-language domains. Our approach emplo

CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration

Model ReleasesDGX agent

arXiv:2605.11186v1 Announce Type: new Abstract: Auto-regressive decoding in Large Language Models (LLMs) is inherently memory-bound: every generation step requires loading the model weights and interm

Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation

Model ReleasesDGX agent

arXiv:2605.11533v1 Announce Type: new Abstract: Clinical check-up reports are multimodal documents that combine page layouts, tables, numerical biomarkers, abnormality flags, imaging findings, and dom

Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters

Model ReleasesDGX agent

arXiv:2605.11960v1 Announce Type: new Abstract: Vision Large Language Models (VLLMs) have achieved remarkable success in modern text-rich visual understanding. However, their perceptual robustness in

clarifying how much “task horizon” falls off as a function of the increasing accuracy criterion directly in the graph (not across tabs) woul…

Model ReleasesDGX agent

clarifying how much “task horizon” falls off as a function of the increasing accuracy criterion directly in the graph (not across tabs) would definitely improve this graph, @METR_Evals. as a variant o

Claude Code and @suno have more in common than you might think: 'It's fun to build things, and it's fun to use what you build.' AI lets peop…

Model ReleasesDGX agent

Claude Code and @suno have more in common than you might think: 'It's fun to build things, and it's fun to use what you build.' AI lets people be creative in almost any domain, from coding to making m

Claude Code weekly limits are increasing 50%, now through July 13. Live now for all Pro, Max, Team, and seat-based Enterprise users.

Model ReleasesDGX agent

Claude Code weekly usage limits have increased by 50% for all Pro, Max, Team, and seat-based Enterprise users, effective immediately through July 13. This temporary expansion applies across all user t

ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IV

Model ReleasesDGX agent

arXiv:2605.11143v1 Announce Type: new Abstract: Reasoning benchmarks measure clinical performance on clean inputs. We evaluate the step before reasoning: retrieval over real EHR notes, where negation,

Control of Fully Actuated Aerial Vehicles: A Comparison of Model-based and Sensor-based Dynamic Inversion

Model ReleasesDGX agent

arXiv:2605.12071v1 Announce Type: new Abstract: Fully actuated multirotor platforms decouple translational force generation from vehicle attitude, enabling independent control of position and orientat

CORE: Cyclic Orthotope Relation Embedding for Knowledge Graph Completion

Model ReleasesDGX agent

arXiv:2605.11159v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to automatically infer missing facts in multi-relational data by mapping entities and relations into continuous re

Correcting Selection Bias in Sparse User Feedback for Large Language Model Quality Estimation: A Multi-Agent Hierarchical Bayesian Approach

Model ReleasesDGX agent

arXiv:2605.12177v1 Announce Type: new Abstract: [Abridged] Production LLM deployments receive feedback from a non-random fraction of users: thumbs sit mostly in the tails of the satisfaction distribut

Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification

Model ReleasesDGX agent

arXiv:2603.28488v2 Announce Type: replace Abstract: Large language models (LLMs) remain unreliable for high-stakes claim verification due to hallucinations and shallow reasoning. While retrieval-augme

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

Model ReleasesDGX agent

arXiv:2605.12501v1 Announce Type: new Abstract: Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions i

Crash Assessment via Mesh-Based Graph Neural Networks and Physics-Aware Attention

Model ReleasesDGX agent

arXiv:2605.11784v1 Announce Type: cross Abstract: Full-vehicle crash simulations are computationally expensive, limiting their use in iterative design exploration. This work investigates learned hybri

CSP Allow-list Experiment

Model ReleasesDGX agent

Tool: CSP Allow-list Experiment An experiment that shows that you can load an app in a CSP-protected sandboxed iframe (see previous note) and have a custom fetch() that intercepts CSP errors and passe

CTFusion: A CTF-based Benchmark for LLM Agent Evaluation

Model ReleasesDGX agent

arXiv:2605.11504v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have enabled agentic systems for complex, multi-step tasks; cybersecurity is emerging as a prominent app

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes

Model ReleasesDGX agent

arXiv:2512.24985v4 Announce Type: replace Abstract: Vision Language Models (VLMs) are increasingly adopted as central reasoning modules for embodied agents. Existing benchmarks evaluate their capabili

Decomposing Evolutionary Mixture-of-LoRA Architectures: The Routing Lever, the Lifecycle Penalty, and a Substrate-Conditional Boundary

Model ReleasesDGX agent

arXiv:2605.11153v1 Announce Type: new Abstract: We decompose an evolutionary mixture-of-LoRA system on a from-scratch ~150M-parameter widened-D substrate (D=1536, V=32000; D/V approx 0.048; the 'widen

Delightful Gradients Accelerate Corner Escape

Model ReleasesDGX agent

arXiv:2605.11908v1 Announce Type: new Abstract: Softmax policy gradient converges at O(1/t), but its transient behavior near sub-optimal corners of the simplex can be exponentially slow. The bottlenec

Deploying Self-Supervised Learning for Real Seismic Data Denoising

Model ReleasesDGX agent

arXiv:2605.11109v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has emerged as a promising approach to seismic data denoising as it does not require clean reference data. In this work

Detecting Data Contamination in LLMs via In-Context Learning

Model ReleasesDGX agent

arXiv:2510.27055v2 Announce Type: replace Abstract: We present Contamination Detection via Context (CoDeC), a practical and accurate method to detect and quantify training data contamination in large

DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows

Model ReleasesDGX agent

arXiv:2304.09479v5 Announce Type: replace Abstract: We introduce a novel approach to single-view face relighting in the wild, addressing challenges such as global illumination and cast shadows. A comm

DiffScore: Text Evaluation Beyond Autoregressive Likelihood

Model ReleasesDGX agent

arXiv:2605.11601v1 Announce Type: new Abstract: Autoregressive language models are widely used for text evaluation, however, their left-to-right factorization introduces positional bias, i.e., early t

DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism

Model ReleasesDGX agent

arXiv:2605.11005v1 Announce Type: new Abstract: Mixture-of-experts (MoE) architectures enable trillion-parameter LLMs with sparsely activated experts. Expert parallelism (EP) is a widely adopted MoE t

← Previous
1…254255256257258…377
Next →