AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
28 May 2026

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding

Model ReleasesDGX agent

arXiv:2505.21771v2 Announce Type: replace-cross Abstract: Multimodal tables i.e. tabular layouts interleaved with charts, maps, icons, and color encodings are ubiquitous in real applications yet remai

Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities

SafetyDGX agent

arXiv:2605.27388v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly utilized as proxies for computational social analysis; yet, their ability to faithfully represent the 't

Modeling Vehicle-Type-Specific Pedestrian Crash Avoidance Behavior in Safety-Critical Interactions Using Smooth-Mamba Deep Reinforcement Learning


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

arXiv:2605.28552v1 Announce Type: new Abstract: As automated vehicles (AVs) increasingly share roadways with human-driven vehicles (HDVs), understanding how pedestrians respond to different vehicle ty

Models That Know How Evaluations Are Designed Score Safer

Model ReleasesDGX agent

arXiv:2605.28591v1 Announce Type: cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified tes

MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents

Model ReleasesDGX agent

arXiv:2605.27853v1 Announce Type: new Abstract: We present MolLingo, a multi-agent system that emulates the reasoning process of a chemist to automate molecular design. Existing LLM-based approaches e

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation

Model ReleasesDGX agent

arXiv:2605.28035v1 Announce Type: new Abstract: In recent years, Multi-Talker Audio-Video Generation (MTAVG) models have shown promising performance on fundamental metrics such as lip-sync and audio-v

Multi-Adapter Representation Interventions via Energy Calibration

Model ReleasesDGX agent

arXiv:2605.28722v1 Announce Type: new Abstract: Representation intervention has emerged as a promising paradigm for aligning large language models toward desired behaviors without modifying model weig

Multi-Agent LLM-based Metamorphic Testing for REST APIs

AgentsDGX agent

arXiv:2605.28321v1 Announce Type: cross Abstract: As REST APIs become an increasingly significant part of software systems, their validation is becoming more critical. Hence, testing and uncovering un

Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors

ApplicationsDGX agent

arXiv:2605.27967v1 Announce Type: cross Abstract: Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), inclu

MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation

Model ReleasesDGX agent

arXiv:2605.28579v1 Announce Type: new Abstract: Large language models (LLMs) have recently advanced text-driven 3D generation, yet Text-to-CAD remains far from supporting industrial product design. Ex

NCSAM Noise-Compensated Sharpness-Aware Minimization for Noisy Label Learning

SafetyDGX agent

arXiv:2601.19947v2 Announce Type: replace-cross Abstract: Learning from Noisy Labels (LNL) remains a fundamental challenge in deep learning because real-world datasets often contain corrupted annotati

Noise Scheduling as Information-Guided Allocation in Diffusion Training

ResearchDGX agent

arXiv:2602.18647v2 Announce Type: replace-cross Abstract: We introduce InfoNoise, an online adaptive noise schedule for diffusion training that reallocates optimization effort toward noise levels wher

Not All NVFP4 QAT Recipes Are Equal: How Architecture and Scale Shape Model Quality for Anomaly Segmentation

ResearchDGX agent

arXiv:2605.27616v1 Announce Type: cross Abstract: Real-time anomaly segmentation demands both high recall and efficient low-precision inference. We study the three-way interaction of model architectur

Not All Pixels Are Equal: Pixel-wise Meta-Learning for Medical Segmentation with Noisy Labels

Model ReleasesDGX agent

arXiv:2511.18894v5 Announce Type: replace-cross Abstract: Medical image segmentation is crucial for clinical applications, but it is frequently disrupted by noisy annotations and ambiguous anatomical

Object-Centric Vision Token Pruning for Vision Language Models

ResearchDGX agent

arXiv:2511.20439v2 Announce Type: replace-cross Abstract: In Vision Language Models (VLMs), vision tokens are quantity-heavy yet information-dispersed compared with language tokens, thus consume too m

OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Interactive Buildings

Model ReleasesDGX agent

arXiv:2605.28168v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated promising capability in generating reward functions for deep reinforcement learning (DRL)-based building

Ocean4Rec: Offline LLM-Derived OCEAN Profiles for Request-Time VOD Reranking

ResearchDGX agent

arXiv:2605.27429v1 Announce Type: cross Abstract: Industrial video-on-demand (VOD) recommenders need richer content understanding, but LLM-as-reranker designs repeat prompt construction, token generat

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning

SafetyDGX agent

arXiv:2604.18530v2 Announce Type: replace Abstract: Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet m

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration

Local AiDGX agent

arXiv:2605.28805v1 Announce Type: cross Abstract: Visual outcomes are increasingly central to multimodal large language models, making reliable and fine-grained verification essential for scaling gene

On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation

ResearchDGX agent

arXiv:2601.06329v2 Announce Type: replace-cross Abstract: Generative spoken language models pretrained on large-scale raw audio can continue a speech prompt with appropriate content while preserving a

On the Intrinsic Limits of Transformer Image Embeddings in Non-Solvable Spatial Reasoning

Model ReleasesDGX agent

arXiv:2601.03048v2 Announce Type: replace-cross Abstract: Vision Transformers (ViTs) excel in semantic recognition but exhibit systematic failures in spatial reasoning tasks such as mental rotation. W

On the Learnability of Test-Time Adaptation: A Recovery Complexity Perspective

ResearchDGX agent

arXiv:2605.28057v1 Announce Type: cross Abstract: Test-time adaptation (TTA) aims to adapt models to maintain reliable performance on non-stationary test streams without requiring labeled data. Despit

On the Origin of Synthetic Information by Means of Steganographic Inheritance

ResearchDGX agent

arXiv:2605.27551v1 Announce Type: new Abstract: The origin of species has been the mystery of mysteries in natural science. By analogy, the origin of synthetic information, we suggest, is the mystery

On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note

Model ReleasesDGX agent

arXiv:2605.27563v1 Announce Type: cross Abstract: This short note presents a dimension-independent subgaussian concentration bound for Gaussian vectors under coordinate-wise nonlinear mappings. Discov

Online Irregular Multivariate Time Series Forecasting via Uncertainty-Driven Dual-Expert Calibration

ApplicationsDGX agent

arXiv:2605.28603v1 Announce Type: cross Abstract: Irregular multivariate time series forecasting is critical in many real-world applications, where time series are irregularly sampled and exhibit dyna

OpenURMA: A Clean-Room Open Implementation of the Unified Bus Protocol

HardwareDGX agent

arXiv:2605.28717v1 Announce Type: new Abstract: Modern datacenter RDMA is bottlenecked at the network interface, not the wire. A NIC running RoCE or InfiniBand holds per-connection state for every (ap

Operational AI Deployment Assurance: Governance-State Orchestration Under Threshold-Sensitive Deployment Conditions -- A Governance Framework for High-Stakes AI Systems

SafetyDGX agent

arXiv:2605.27827v1 Announce Type: new Abstract: AI governance frameworks increasingly emphasize fairness, transparency, accountability, and lifecycle risk management in high-stakes domains. However, m

Optimal and Diffusion Transports in Machine Learning

ResearchDGX agent

arXiv:2512.06797v2 Announce Type: replace-cross Abstract: Several problems in machine learning are naturally expressed as the design and analysis of time-evolving probability distributions. This inclu

Optimal LTLf Synthesis

Model ReleasesDGX agent

arXiv:2605.11544v2 Announce Type: replace Abstract: Strategy synthesis typically follows an all-or-nothing paradigm, returning unrealisable whenever a specification cannot be guaranteed in an uncertai

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents

Model ReleasesDGX agent

arXiv:2605.28158v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used to assist with operations research (OR) modeling, yet existing OR-oriented benchmarks often redu

Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline

Model ReleasesDGX agent

arXiv:2605.27440v1 Announce Type: cross Abstract: Small changes to how a buyer phrases a question -- 'best CRM' vs 'top CRM' vs 'best CRM for a SaaS startup' -- produce substantially different brand r

Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN

ResearchDGX agent

arXiv:2506.10138v3 Announce Type: replace-cross Abstract: We partially reverse-engineer a convolutional recurrent neural network (RNN) trained with model-free reinforcement learning to play the box-pu

PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft

Model ReleasesDGX agent

arXiv:2605.27762v1 Announce Type: new Abstract: We present PEAM, a Parametric Embodied Agent Memory framework in Minecraft that transforms agent memory from inference-time retrieval into parameter-res

Performance and Explainability Requirements of Evolutionary Algorithms in Real-World Physics-Informed Optimization

ApplicationsDGX agent

arXiv:2605.28164v1 Announce Type: cross Abstract: Evolutionary computation offers a variety of tools to solve complex real-world optimization problems. However, research often focuses on smaller, simp

Periodic RoPE for Infinite Context LLMs

Model ReleasesDGX agent

arXiv:2605.27980v1 Announce Type: cross Abstract: The ability to process ultra-long contexts is crucial for large language models (LLMs) to perform long-horizon tasks. While recent efforts have extend

Personalized Observation Normalization for Federated Reinforcement Learning in Simulation Environments with Heterogeneity

Model ReleasesDGX agent

arXiv:2605.27385v1 Announce Type: cross Abstract: Federated reinforcement learning (FedRL) enables multiple agents to collaboratively train a global policy without sharing raw data, making it ideal fo

Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models

Model ReleasesDGX agent

arXiv:2503.01829v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) demonstrate persuasive capabilities that rival human-level persuasion. While these capabilities can be used for s

PetroBench: A Benchmark for Large Language Models in Petroleum Engineering

Model ReleasesDGX agent

arXiv:2605.28032v1 Announce Type: new Abstract: Large Language Models are increasingly applied in the petroleum industry, highlighting the need for a domain-specific evaluation framework. This study d

Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains

SafetyDGX agent

arXiv:2605.28345v1 Announce Type: new Abstract: Progress in Prognostics and Health Management (PHM) is hindered by the lack of standardized and reusable evaluation practices across tasks, datasets, an

PIRS: Physics-Informed Reward Shaping for SAC-Based Building Energy Management

AgentsDGX agent

arXiv:2605.28232v1 Announce Type: new Abstract: Occupant comfort and grid-aware energy efficiency are competing objectives whose joint optimization depends critically on how reward functions are speci

Plan Before Search: Search Agents Need Plan

AgentsDGX agent

arXiv:2605.28354v1 Announce Type: new Abstract: Training large language models as retrieval-augmented reasoning agents typically combines reinforcement learning with an SFT cold start distilled from a

Planning a Community Approach to Diabetes Care in Low- and Middle-Income Countries Using Optimization

ResearchDGX agent

arXiv:2305.06426v2 Announce Type: replace Abstract: Diabetes is a global health priority, especially in low- and-middle-income countries, where over 50% of premature deaths are attributed to high bloo

Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

Model ReleasesDGX agent

arXiv:2605.28201v1 Announce Type: new Abstract: Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into ext

PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

Model ReleasesDGX agent

arXiv:2605.27887v1 Announce Type: new Abstract: LLMs have shown strong performance across diverse financial tasks, yet portfolio management (PM), a critical financial decision-making task, remains poo

Position: Retire the 'Positive Backdoor' Label -- Secret Alignment Requires Strict and Systematic Evaluation

SafetyDGX agent

arXiv:2605.28597v1 Announce Type: cross Abstract: This position paper argues that the AI/ML community should stop overclaiming and retire the label 'positive backdoor,' and instead treat trigger-activ

Preference-Shaped Expected Hypervolume and R2 Improvement: Exact Computation and Monotonicity

ResearchDGX agent

arXiv:2605.28746v1 Announce Type: cross Abstract: This paper studies preference-shaped expected improvement criteria for Bayesian multiobjective optimization. We consider two indicator families which

Prefix-Safe Bayesian Belief Tracking for LLM Reasoning Reliability:Separating Calibration from Ranking

ResearchDGX agent

arXiv:2605.27712v1 Announce Type: new Abstract: Long reasoning traces need reliability estimates before final answers are known. We study prefix-conditioned eventual-success estimation, P(y=1 mid o_{1

Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

Model ReleasesDGX agent

arXiv:2605.27958v1 Announce Type: cross Abstract: Linear probes trained on LLM activations are increasingly proposed as deception-detection metrics, yet report AUROC exceeding 0.96 on clean benchmarks

Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

SafetyDGX agent

arXiv:2602.01745v2 Announce Type: replace-cross Abstract: Token-level reweighting is a simple yet effective mechanism for controlling supervised fine-tuning, but common indicators are largely one-dime

Probing for Knowledge Attribution in Large Language Models

Model ReleasesDGX agent

arXiv:2602.22787v2 Announce Type: replace-cross Abstract: Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, w

Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit

Model ReleasesDGX agent

arXiv:2605.27439v1 Announce Type: cross Abstract: AI assistants like ChatGPT and Claude are recommendation engines, not search engines: they answer commercial queries by directly nominating brands rat

Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement

Model ReleasesDGX agent

arXiv:2605.28360v1 Announce Type: new Abstract: Automatic prompt optimization (APO) has driven significant gains in LLM-based agentic workflows. However, existing methods treat each task's prompt as a

PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting

Model ReleasesDGX agent

arXiv:2605.28066v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated remarkable efficacy in text embedding, yet current adaptation methods like LoRA face significant bottle

ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation

SafetyDGX agent

arXiv:2605.28293v1 Announce Type: cross Abstract: Proactive Recommender Systems (PRSs) aim to guide user preference shift toward target items by generating paths of intermediate recommendations. Reinf

ProvMind: Provenance-grounded reasoning for materials synthesis

Model ReleasesDGX agent

arXiv:2605.28487v1 Announce Type: new Abstract: Materials process optimization requires reasoning over routes, conditions, tools and causal dependencies, yet most computational formulations flatten sy

PrunePath: Towards Highly Structured Sparse Language Models

Model ReleasesDGX agent

arXiv:2605.28283v1 Announce Type: cross Abstract: Feed-forward networks (FFNs) dominate the parameter count and computation of modern language models, yet existing pruning methods often struggle to co

Pruning and Distilling Mixture-of-Experts into Dense Language Models

Model ReleasesDGX agent

arXiv:2605.28207v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory,

Quantum Machine Learning-based 6G edge Network: Enabling Adaptive Communication and Model Aggregation

ResearchDGX agent

arXiv:2605.27417v1 Announce Type: cross Abstract: With the advent of sixth-generation (6G) mobile communication technology, vehicle-to-everything (V2X) communication faces unprecedented challenges in

QuITE: Query-Based Irregular Time Series Embedding

ApplicationsDGX agent

arXiv:2605.28166v1 Announce Type: cross Abstract: Irregular Multivariate Time Series (IMTS) are common in practice, yet their irregular sampling complicates effective modeling. Existing approaches typ

RAG-Coding: Enhancing LLM Medical Coding with Structured External Knowledge

AgentsDGX agent

arXiv:2605.27377v1 Announce Type: cross Abstract: We present RAG-Coding, an agentic method for automated ICD-10-CM coding. RAG-Coding orchestrates four large language model (LLM) agents and grounds th

← Previous
1…200201202203204…358
Next →