AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,356
  • Agents7,554
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,165
  • Local Ai4,929
  • Model Releases23,869
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,356
  • Agents7,554
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,165
  • Local Ai4,929
  • Model Releases23,869
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

88,356Total entries
1Added by human
88,355Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,584 results
29 May 2026

Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering

Model ReleasesDGX agent

arXiv:2605.29648v1 Announce Type: new Abstract: Applying reinforcement learning to improve factual accuracy in knowledge-intensive question answering faces a reward design dilemma. Response-level rewa

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation

Model ReleasesDGX agent

arXiv:2605.30317v1 Announce Type: new Abstract: Autoregressive image and video generators are trained with teacher-forced histories but must sample from their own generated prefixes at inference time,

28 May 2026

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2605.28556v1 Announce Type: new Abstract: As agent capabilities advance, existing benchmarks, such as au^2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remain

AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens

Model ReleasesDGX agent

arXiv:2512.17375v2 Announce Type: replace-cross Abstract: LLM-as-a-Judge systems supply the reward signal in modern RLHF and RLVR pipelines, but their binary verdict reduces to a single linear readout

AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems

SafetyDGX agent

arXiv:2605.27466v1 Announce Type: cross Abstract: Multi-agent systems built on large language models (LLMs) require many coordination choices that are difficult to fix a priori: which skill protocol t

Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers

Model ReleasesDGX agent

arXiv:2512.09800v2 Announce Type: replace Abstract: Low-power microcontroller (MCU) hardware is currently evolving from single-core architectures to predominantly multi-core architectures. In parallel

ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference

Model ReleasesDGX agent

arXiv:2505.19342v2 Announce Type: replace-cross Abstract: Multi-device inference can reduce Transformer latency by parallelizing computation. However, existing methods require high inter-device bandwi

Asynchronous Remote Sensing Time-Series Fusion for Cloud Removal and Anytime Reconstruction

Model ReleasesDGX agent

arXiv:2605.27726v1 Announce Type: new Abstract: Frequent cloud cover severely limits the usability of Sentinel-2 (S2) optical time series for Earth surface monitoring. Sentinel-1 (S1) SAR provides all

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

Model ReleasesDGX agent

arXiv:2605.28508v1 Announce Type: new Abstract: Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape us

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week @every and our verdict is they could've ju…

Model ReleasesDGX agent

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week @every and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check:

Building Community-Centred NLP Resources for Puno Quechua

Model ReleasesDGX agent

arXiv:2605.28253v1 Announce Type: new Abstract: The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR

Comparative Analysis of Liquid Neural Networks and LSTM for Sequential Pattern Recognition: Robustness, Efficiency, and Clinical Utility

Model ReleasesDGX agent

arXiv:2605.27467v1 Announce Type: cross Abstract: Traditional Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) units operate on discrete time steps, often failing to capture the flui

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization

SafetyDGX agent

arXiv:2605.28615v1 Announce Type: new Abstract: Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex compositional prompts (covering attribute bi

Decision-focused learning for optimal PV-Battery scheduling

ResearchDGX agent

arXiv:2605.28340v1 Announce Type: cross Abstract: The use of residential photovoltaics has increased dramatically in recent years. With battery systems becoming more affordable, the optimal operation

Detection Without Correction: A Two-Parameter Decomposition of Multi-Stage LLM Pipelines

Model ReleasesDGX agent

arXiv:2605.27559v1 Announce Type: cross Abstract: Multi-stage LLM pipelines that perform multi-agent debate, intrinsic self-correction, or retrieval-augmented verification exhibit puzzling aggregate b

DisasterBench: Benchmarking LLM Planning under Typed Tool Interface Constraints

Model ReleasesDGX agent

arXiv:2605.27957v1 Announce Type: new Abstract: Disasters cause severe societal impacts, demanding rapid coordination of heterogeneous AI tools, from satellite analysis to flood prediction and damage

Disentangling Language Roles in Multilingual LLM Task Execution

Model ReleasesDGX agent

arXiv:2605.27649v1 Announce Type: new Abstract: Multilingual LLMs are increasingly used when instruction, source content, and required response languages do not coincide. Existing benchmarks have expa

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning

Model ReleasesDGX agent

arXiv:2510.27266v2 Announce Type: replace Abstract: Autonomous graphical user interface (GUI) agents rely on accurate GUI grounding, which maps language instructions to on-screen coordinates, to execu

FedEHR-Gen: Federated Synthetic Time-Series EHR Generation via Latent Space Alignment and Distribution-Aware Aggregation

SafetyDGX agent

arXiv:2605.27892v1 Announce Type: new Abstract: Synthetic Electronic Health Record (EHR) generation provides a promising avenue for data augmentation and cross-hospital modeling in privacy-constrained

From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence

Model ReleasesDGX agent

arXiv:2605.28371v1 Announce Type: new Abstract: Industrial Prognostics and Health Management (PHM) provides a representative case study for a broader challenge in applied machine learning: translating

HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains

Model ReleasesDGX agent

arXiv:2605.28315v1 Announce Type: new Abstract: General-purpose machine translation benchmarks such as FLORES-200 have reached a saturation regime on Chinese-English pairs, where modern large language

Integrated and Cross-Architecture Interpretation of LLM Reasoning

Model ReleasesDGX agent

arXiv:2605.28006v1 Announce Type: cross Abstract: Understanding how LLMs reason is hindered by a practical asymmetry: while their generated outputs are observable, the underlying reasoning patterns re

IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents

Model ReleasesDGX agent

arXiv:2605.28714v1 Announce Type: cross Abstract: An Initial Public Offering (IPO) filing is a document released when a private firm goes public, allowing individual (retail) investors to purchase its

Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking

Model ReleasesDGX agent

arXiv:2605.27914v1 Announce Type: cross Abstract: Subjective evaluation of LLM behavior -- empathy, restraint, calibrated emotional tone -- is hard. Human inter-rater agreement on such qualities satur

LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

Model ReleasesDGX agent

arXiv:2605.28721v1 Announce Type: new Abstract: Are LLM-based search agents genuinely searching, or using the web to verify what they already know? We study this question on BrowseComp with three diag

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

Model ReleasesDGX agent

arXiv:2603.21165v2 Announce Type: replace Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresente

MetaboT: An LLM-based Multi-Agent Frameworkfor Interactive Analysis of Mass SpectrometryMetabolomics Knowledge Graphs

Model ReleasesDGX agent

arXiv:2510.01724v2 Announce Type: replace Abstract: Mass spectrometry-based metabolomics generates complex, high-dimensional data that holds vast potential for biological discovery but remains difficu

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding

Model ReleasesDGX agent

arXiv:2505.21771v2 Announce Type: replace-cross Abstract: Multimodal tables i.e. tabular layouts interleaved with charts, maps, icons, and color encodings are ubiquitous in real applications yet remai

MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents

Model ReleasesDGX agent

arXiv:2605.27853v1 Announce Type: new Abstract: We present MolLingo, a multi-agent system that emulates the reasoning process of a chemist to automate molecular design. Existing LLM-based approaches e

No Safe Dose: How Training Data Drives Unsafe Image Generation

SafetyDGX agent

arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remai

Not All Pixels Are Equal: Pixel-wise Meta-Learning for Medical Segmentation with Noisy Labels

Model ReleasesDGX agent

arXiv:2511.18894v5 Announce Type: replace-cross Abstract: Medical image segmentation is crucial for clinical applications, but it is frequently disrupted by noisy annotations and ambiguous anatomical

On the Equivariant Learning of the Q-tensor Order Parameter

Model ReleasesDGX agent

arXiv:2605.27679v1 Announce Type: cross Abstract: We construct and evaluate group-equivariant neural networks for the prediction of the two-dimensional Q-tensor order parameter of nematic liquid cryst

Opus 4.8 is now supported in Hermes Agent ^_^

AgentsDGX agent

Nous Research has announced support for Opus 4.8 in the Hermes Agent framework. This update enables the Hermes Agent to utilize Anthropic's Opus 4.8 model, expanding the available model options for us

Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline

Model ReleasesDGX agent

arXiv:2605.27440v1 Announce Type: cross Abstract: Small changes to how a buyer phrases a question -- 'best CRM' vs 'top CRM' vs 'best CRM for a SaaS startup' -- produce substantially different brand r

PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

Model ReleasesDGX agent

arXiv:2605.27887v1 Announce Type: new Abstract: LLMs have shown strong performance across diverse financial tasks, yet portfolio management (PM), a critical financial decision-making task, remains poo

PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature

Model ReleasesDGX agent

arXiv:2605.28375v1 Announce Type: new Abstract: Prion diseases are rare, rapidly progressive, and fatal neurodegenerative disorders that remain difficult to diagnose, particularly in their early stage

Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

SafetyDGX agent

arXiv:2510.06974v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they may reflect and amplify social biases

Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit

Model ReleasesDGX agent

arXiv:2605.27439v1 Announce Type: cross Abstract: AI assistants like ChatGPT and Claude are recommendation engines, not search engines: they answer commercial queries by directly nominating brands rat

PubMedCausal: A Span-Level Annotated Corpus for Causal Relation Extraction in Biomedical Text

Model ReleasesDGX agent

arXiv:2605.28363v1 Announce Type: new Abstract: Causal relation extraction (CRE) is central to biomedical text mining, but current resources often conflate causal relations with broader associations,

ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing

Model ReleasesDGX agent

arXiv:2511.14584v3 Announce Type: replace-cross Abstract: We present ReflexGrad, a dual-process architecture for within-episode failure recovery in LLM agents without demonstrations. When agents commi

Relational Semantic Reasoning on 3D Scene Graphs for Open World Interactive Object Search

Model ReleasesDGX agent

arXiv:2603.05642v2 Announce Type: replace-cross Abstract: Open-world interactive object search in household environments requires understanding semantic relationships between objects and their surroun

Rethinking Visual Neglect: Steering via Context-Preference for MLLM Hallucination Mitigation

ResearchDGX agent

arXiv:2605.27993v1 Announce Type: new Abstract: Object hallucination remains a primary obstacle to the reliable deployment of Multimodal Large Language Models (MLLMs). Current inference-time mitigatio

RMPL: Relation-aware Multi-task Progressive Learning with Stage-wise Training for Multimedia Event Extraction

Model ReleasesDGX agent

arXiv:2602.13748v2 Announce Type: replace Abstract: Multimedia Event Extraction (MEE) aims to identify events and their arguments from documents that contain both text and images. It requires groundin

SAM-Enhanced Segmentation on Road Datasets: Balancing Critical Classes in Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.28136v1 Announce Type: new Abstract: Dense semantic segmentation is essential for autonomous driving, yet many multi-modal datasets lack pixel-level annotations. The Zenseact Open Dataset (

Self-Prophetic Decoding to Unlock Visual Search in LVLMs

ResearchDGX agent

arXiv:2605.28741v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) are rapidly evolving toward true multimodal reasoning, with visual search representing a concrete instantiation of

StoryLens: Preference-Aligned Story Rewriting via Context-Aware Narrative Enrichment

Model ReleasesDGX agent

arXiv:2605.28073v1 Announce Type: cross Abstract: Story rewriting aims to adapt existing narratives to diverse reader preferences while preserving plot consistency and narrative coherence. Unlike conv

Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey

Model ReleasesDGX agent

arXiv:2605.27431v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across dive

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

Model ReleasesDGX agent

arXiv:2605.28063v1 Announce Type: cross Abstract: Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. C

Unsupervised Identification and Removal of Spurious Correlations During Fine-Tuning

SafetyDGX agent

arXiv:2605.27676v1 Announce Type: cross Abstract: Fine-tuning a pretrained language model on a curated dataset can produce spurious correlations between the fine-tuning task and unintended latent fact

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

Model ReleasesDGX agent

arXiv:2605.27882v1 Announce Type: cross Abstract: LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning

Model ReleasesDGX agent

arXiv:2510.08555v2 Announce Type: replace Abstract: Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpaint

Weak Convergence Analysis of Online Neural Actor-Critic Algorithms

Model ReleasesDGX agent

arXiv:2403.16825v2 Announce Type: replace Abstract: We prove that a single-layer neural network trained with the online actor critic algorithm converges in distribution to a random ordinary differenti

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

SafetyDGX agent

arXiv:2605.27932v1 Announce Type: cross Abstract: Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly underst

You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

Model ReleasesDGX agent

arXiv:2605.27586v1 Announce Type: cross Abstract: Ensuring agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. W

27 May 2026

COVD: Continual Open-Vocabulary Object Detection with Novel Concept Injection

Model ReleasesDGX agent

arXiv:2605.27116v1 Announce Type: new Abstract: Open-vocabulary object detection (OVD) has made significant progress, enabling detectors to generalize from seen to unseen categories. However, real-wor

Curation and Extraction of Drug-Related Entities from Reddit Platform

Model ReleasesDGX agent

arXiv:2605.26445v1 Announce Type: new Abstract: Physicians learn primarily about illicit drugs from clinical overdose cases, limiting their understanding of real-world usage. Meanwhile, drug users sha

DGLD: Domain-Gated Latent Diffusion for the Discovery of Novel Energetic Materials

Model ReleasesDGX agent

arXiv:2605.26540v1 Announce Type: cross Abstract: Energetic-materials performance gains translate directly into reduced propellant mass, smaller warheads, and more efficient civilian gas-generators, y

Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?

ApplicationsDGX agent

arXiv:2605.27135v1 Announce Type: cross Abstract: With the rapid proliferation of generative models, such as diffusion models, digital watermarking has emerged as a crucial solution for identifying AI

FedTreeLoRA: Reconciling Statistical and Functional Heterogeneity in Federated LoRA Fine-Tuning

Model ReleasesDGX agent

arXiv:2603.13282v2 Announce Type: replace-cross Abstract: Federated Learning (FL) with Low-Rank Adaptation (LoRA) has become a standard for privacy-preserving LLM fine-tuning. However, existing person

Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards

Model ReleasesDGX agent

arXiv:2605.26579v1 Announce Type: new Abstract: The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement lea

← Previous
1…467468469470471…1060
Next →