AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

88,419Total entries
1Added by human
88,418Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,638 results
27 May 2026

SteelDS: A High-Resolution Video Dataset of E40 Steel Scrap for Object Detection and Instance Segmentation

Model ReleasesDGX agent

arXiv:2605.26682v1 Announce Type: cross Abstract: This dataset provides high-resolution, annotated video sequences of shredded E40-grade steel and copper scrap on a conveyor belt. Captured in a contro

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

SafetyDGX agent

arXiv:2601.05729v2 Announce Type: replace Abstract: Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for t

Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.27239v1 Announce Type: new Abstract: Annotation quality is difficult to sustain when campaigns span weeks or months with small annotator pools. We present a Setswana sentiment dataset of 3,

Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records

Model ReleasesDGX agent

arXiv:2605.26463v1 Announce Type: cross Abstract: Data consistency between unstructured clinical notes and structured tables in Electronic Health Records (EHRs) is essential for patient safety and cli

Towards Interpretable Federated Learning

Local AiDGX agent

arXiv:2302.13473v2 Announce Type: replace Abstract: Federated learning (FL) enables multiple data owners to build machine learning models collaboratively without exposing their private local data. In

Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial VOCs in the Steel Industry

Model ReleasesDGX agent

arXiv:2605.27071v1 Announce Type: new Abstract: Key knowledge for steel-industry volatile organic compounds (VOCs) governance is scattered across unstructured scientific literature, making it difficul

Trust Region Q Adjoint Matching

Model ReleasesDGX agent

arXiv:2605.27079v1 Announce Type: cross Abstract: Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step s

UCPO: Uncertainty-Aware Policy Optimization

SafetyDGX agent

arXiv:2601.22648v2 Announce Type: replace Abstract: The key to building trustworthy large language models (LLMs) lies in endowing them with inherent uncertainty expression capabilities, thereby mitiga

Uncertainty-Aware Budget Allocation for Adaptive Test-Time Reasoning

ResearchDGX agent

arXiv:2605.26849v1 Announce Type: new Abstract: Sampling multiple responses improves language model reasoning, but uniform compute allocation is inefficient: easy questions are over-sampled while hard

Underwater360: Reconstructing Underwater Scenes from Panoramic Images with Omnidirectional Gaussian Splatting

Model ReleasesDGX agent

arXiv:2605.26447v1 Announce Type: new Abstract: Underwater scene reconstruction is essential for immersive exploration of aquatic environments, yet remains challenging due to complex participating-med

Variational Inference for Evidential Deep Learning

Model ReleasesDGX agent

arXiv:2605.26477v1 Announce Type: new Abstract: While Deep Neural Networks (DNNs) achieve remarkable performance, their tendency to produce overconfident predictions. Evidential Deep Learning (EDL) mi

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

Model ReleasesDGX agent

arXiv:2605.26144v1 Announce Type: cross Abstract: We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents. Unlike

What Molecular Structure Cannot Tell Us: A Taxonomy of Explainability Gaps in GNN-Based Drug Toxicity Prediction

Model ReleasesDGX agent

arXiv:2605.26183v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) have emerged as a structurally natural approach for molecular toxicity prediction, operating directly on atomic connectiv

Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion

Model ReleasesDGX agent

arXiv:2605.26383v1 Announce Type: new Abstract: Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and l

26 May 2026

A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2605.25502v1 Announce Type: cross Abstract: Educational aspect-based sentiment analysis (ABSA) can support course improvement, but public aspect-labeled student feedback remains scarce because e

A lift for input-convex neural network training

Model ReleasesDGX agent

arXiv:2605.24274v1 Announce Type: new Abstract: Input-convex neural networks (ICNNs) are widely used for log-concave density estimation, convex-potential normalizing flows, optimal transport, and tran

A Lightweight Hybrid Transformer-CRF Architecture for Multi-Type Bangla Medical Entity Recognition

ApplicationsDGX agent

arXiv:2605.25463v1 Announce Type: new Abstract: MedER refers to the identification of medical entities. It is crucial for extracting structured clinical information from unstructured medical text. Man

A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays

Model ReleasesDGX agent

arXiv:2605.25652v1 Announce Type: new Abstract: Free-form legal essay evaluation in NLP treats expert inter-rater stability as a single ceiling number, and treats LLM-judge agreement with that ceiling

Action-Prior Denoising for Smooth Real-Time Chunking

Model ReleasesDGX agent

arXiv:2605.25537v1 Announce Type: new Abstract: Real-time chunking (RTC) lets chunked action policies operate under inference delay by conditioning a newly generated action chunk on actions already co

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

Model ReleasesDGX agent

arXiv:2605.25272v1 Announce Type: new Abstract: While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, ma

Algometrics: Forecasting Under Algorithmic Feedback

ResearchDGX agent

arXiv:2605.23978v1 Announce Type: new Abstract: In algorithmic markets, predictive models become part of the data-generating process they aim to forecast. Once their outputs are converted into trades,

An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods

TutorialsDGX agent

arXiv:2605.24298v1 Announce Type: cross Abstract: The growing use of Large Language Models (LLMs) for automated code generation has enhanced software development efficiency, but often at the cost of s

ASTRO: Adaptive Spatio-Temporal Reinforcement Optimization for GNN Powered Anomly Detection in Cyber Physical Systems

ApplicationsDGX agent

arXiv:2605.25135v1 Announce Type: cross Abstract: Anomaly detection in Industrial Internet of Things (IIoT) environments is essential to protect the Industrial Control Systems (ICS) and Cyber-Physical

AuthTrace: Diagnosing Evidence Construction in Thematically Dense Single-Author Corpora

Model ReleasesDGX agent

arXiv:2605.25382v1 Announce Type: new Abstract: Evidence construction systems--chunk retrieval, agent memory, knowledge-graph traversal, and thematic indexing--are evaluated on separate benchmarks wit

Benchmarking and Learning Real-World Customer Service Dialogue

Model ReleasesDGX agent

arXiv:2510.22143v3 Announce Type: replace Abstract: Existing benchmarks and training pipelines for industrial intelligent customer service (ICS) remain misaligned with real-world dialogue requirements

Binding Visual Features Point by Point

ResearchDGX agent

arXiv:2605.25427v1 Announce Type: cross Abstract: Despite success on standard benchmarks, vision language models display persistent failures on tasks involving processing of multi-object scenes, inclu

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference

ResearchDGX agent

arXiv:2511.16449v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown great potential for embodied AI by integrating visual perception, language understanding, and a

CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures

AgentsDGX agent

arXiv:2605.25338v1 Announce Type: cross Abstract: Large language model (LLM) agents frequently fail on multi-step tasks involving reasoning, tool use, and environment interaction. While such failures

Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA

Model ReleasesDGX agent

arXiv:2605.25204v1 Announce Type: new Abstract: Pluralistic alignment requires systems to adapt to diverse user values, communication styles, and contextual assumptions. We believe that a foundational

Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation

ResearchDGX agent

arXiv:2605.25831v1 Announce Type: cross Abstract: Large language models (LLMs) define a distribution over text, which can be viewed as a probabilistic representation of uncertainty: sampling K respons

Complement Submodular Information Measures for Balanced and Robust Data Selection

Model ReleasesDGX agent

arXiv:2605.24779v1 Announce Type: cross Abstract: Submodular optimization has become a fundamental paradigm for data selection, retrieval, summarization, and representation learning due to its ability

ConceptM^3oE: Concept-Guided Multimodal Mixture of Experts for Interpretable Computational Pathology

ApplicationsDGX agent

arXiv:2605.24399v1 Announce Type: new Abstract: Healthcare models are transitioning from unimodal prediction toward multimodal reasoning over heterogeneous diagnostic inputs. In computational patholog

Critical Organization of Deep Neural Networks, and p-Adic Statistical Field Theories

Model ReleasesDGX agent

arXiv:2601.19070v2 Announce Type: replace Abstract: We rigorously study the thermodynamic limit of deep neural networks (DNNS) and recurrent neural networks (RNNs), assuming that the activation functi

CSP-Atlas: Concept-Specific Neural Circuits in a Sparse Python Transformer

Model ReleasesDGX agent

arXiv:2605.24603v1 Announce Type: new Abstract: A sparse 8-layer code transformer develops dedicated neural circuitry for every Python construct tested, and that circuitry is organised by a clean comp

CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents

Model ReleasesDGX agent

arXiv:2605.25624v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its exte

CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning

Model ReleasesDGX agent

arXiv:2605.24331v1 Announce Type: new Abstract: Context or prompt-level reweighting has emerged as a central algorithmic lever in Reinforcement Learning with Verified Rewards (RLVR) for improving the

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations

AgentsDGX agent

arXiv:2605.24539v1 Announce Type: new Abstract: Agent harness evolution improves frozen language-model agents by modifying the executable structures around them. We study this paradigm as a form of sa

Deployment-complete benchmarking

Model ReleasesDGX agent

arXiv:2605.25997v1 Announce Type: new Abstract: Benchmarks increasingly guide deployment, procurement and scientific screening, yet a score supports only the response it records, not necessarily the d

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

SafetyDGX agent

arXiv:2605.23975v1 Announce Type: new Abstract: Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Foc

DisDop: Distillation with Domain Priors for Open-Vocabulary Aerial Object Detection

SafetyDGX agent

arXiv:2605.24639v1 Announce Type: cross Abstract: With the widespread application of drones in recent years, object detection of aerial images has attracted increasing attention, especially open-vocab

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning

SafetyDGX agent

arXiv:2605.25604v1 Announce Type: new Abstract: Reinforcement Learning has become a standard paradigm for aligning Large Language Models with human intent and task requirements. While Group Relative P

DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition

Model ReleasesDGX agent

arXiv:2512.11941v2 Announce Type: replace-cross Abstract: Zero-shot skeleton-based action recognition (ZS-SAR) is fundamentally constrained by prevailing approaches that rely on aligning skeleton feat

Efficient Benchmarking Is Just Feature Selection and Multiple Regression

Model ReleasesDGX agent

arXiv:2605.25773v1 Announce Type: cross Abstract: Efficient benchmarking techniques aim to lower the computational cost of evaluating LLMs by predicting full benchmark scores using only a subset of a

EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization

Model ReleasesDGX agent

arXiv:2605.25395v1 Announce Type: new Abstract: Lookahead-based acceleration methods, such as Nesterov's momentum, are widely used in optimization, but they often become unreliable in deep learning tr

Emission-Aware Reinforcement Learning for Sustainable Electric Vehicle Charging and Carbon Dioxide Reduction Under Varying Renewable Penetration

Model ReleasesDGX agent

arXiv:2605.24543v1 Announce Type: new Abstract: The rapid growth of Electric Vehicle (EV) adoption challenges power distribution networks through peak load spikes, voltage instability, and transformer

End-to-End Intracortical Speech Decoding from Neural Activity

ResearchDGX agent

arXiv:2605.24313v1 Announce Type: new Abstract: Current high-performing intracortical speech neuroprostheses achieve low word error rates but typically rely on external language models during inferenc

Equip Pre-ranking with Target Attention by Residual Quantization

ApplicationsDGX agent

arXiv:2509.16931v3 Announce Type: replace-cross Abstract: The pre-ranking stage in industrial recommendation systems faces a fundamental conflict between efficiency and effectiveness. While powerful m

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture

ResearchDGX agent

arXiv:2605.24144v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved impressive performance across diverse domains but remain inefficient during the autoregressive decoding pha

everybody talks about the china->us catchup not enough people talking about the us-> china catchup great job @o_lacombe et al, @robert_mchar…

Model ReleasesDGX agent

everybody talks about the china->us catchup not enough people talking about the us-> china catchup great job @o_lacombe et al, @robert_mchardy et al! [AINews 3 Apr 2026] Gemma 4: The world's best smal

EvoEGF-Mol: Evolving Exponential Geodesic Flow for Structure-based Drug Design

Model ReleasesDGX agent

arXiv:2601.22466v2 Announce Type: replace Abstract: Structure-Based Drug Design (SBDD) aims to discover bioactive ligands. Conventional approaches construct probability paths separately in Euclidean a

Explore Before You Solve: The Speed--Depth Trade-off in Epistemic Agents for ARC-AGI-3

Model ReleasesDGX agent

arXiv:2605.25931v1 Announce Type: new Abstract: We systematically investigate all 25 public ARC-AGI-3 games and find that every one is reachable through non-intelligent strategies: 10 in a single blin

Extending Embodied Question Answering from Perception to Decision

Model ReleasesDGX agent

arXiv:2605.25813v1 Announce Type: new Abstract: Embodied Question Answering (EQA) connects perception, reasoning, and interaction within embodied environments. However, existing datasets and benchmark

Factorize to Generalize: Retrieval-Guided Invariant-Dynamic Decomposition for Time Series Forecasting

TutorialsDGX agent

arXiv:2605.24911v1 Announce Type: cross Abstract: Time series foundation models (TSFMs) have recently achieved strong zero-shot forecasting performance through large-scale pretraining and retrieval-au

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning

SafetyDGX agent

arXiv:2605.24286v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning is useful for monitoring language models only when the reasoning trace faithfully reflects the computation that produ

False Fixed Points: Kantian Feedback, Stable Miscalibration, and Representational Compression in LLMs

Local AiDGX agent

arXiv:2510.14925v4 Announce Type: replace Abstract: High-confidence errors in large language models are often treated as fragile failures. We study an alternative: some errors may be false fixed point

Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers

ResearchDGX agent

arXiv:2603.05143v3 Announce Type: replace Abstract: Understanding reasoning in large language models is complicated by evaluations that conflate multiple reasoning types. We isolate analogical reasoni

FLOATBench: A Dataset and Benchmark for Floating Offshore Wind Turbine Tower Fatigue

Model ReleasesDGX agent

arXiv:2605.25717v1 Announce Type: new Abstract: Most of the world's offshore wind resource lies in waters too deep for fixed-bottom foundations, making floating offshore wind turbines (FOWTs) essentia

From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP

SafetyDGX agent

arXiv:2605.25226v1 Announce Type: new Abstract: Large language models are widely deployed in high-stakes NLP tasks, yet risks such as bias, hallucination, adversarial vulnerability and unreliable gene

From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents

Model ReleasesDGX agent

arXiv:2605.25693v1 Announce Type: new Abstract: While role-playing agents excel in short-term interactions, long-term conversations overwhelm context windows, motivating external memory frameworks. Cu

Future-KL Regularized GRPO: Process-Level Credit Assignment from f-Divergence Regularization

Local AiDGX agent

arXiv:2601.10201v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) is widely used for critic-free Large Language Model (LLM) post-training, but its KL regularization i

← Previous
1…548549550551552…1061
Next →