AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
7 Aug 2026

ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

Model ReleasesDGX agent

arXiv:2608.06110v1 Announce Type: new Abstract: This paper presents ECHO (Enhanced Care & Health Observer), a locally-deployable conversational health assistant for long-term chronic care management.

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

Model ReleasesDGX agent

arXiv:2608.05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local look

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.06197v1 Announce Type: new Abstract: Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose

Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows

SafetyDGX agent

arXiv:2608.05602v1 Announce Type: new Abstract: Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and

Estimating time spent on work tasks

SafetyDGX agent

arXiv:2608.05172v1 Announce Type: cross Abstract: The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affects work: a

Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

Model ReleasesDGX agent

arXiv:2608.05411v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learni

Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents

Model ReleasesDGX agent

arXiv:2608.06108v1 Announce Type: new Abstract: Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, p

Evidential Rule Learning for Interpretable Classification with Abstention

Model ReleasesDGX agent

arXiv:2608.05859v1 Announce Type: cross Abstract: Interpretable classification often requires more than accurate predictions for real-life deployment: models should be transparent about the evidence b

Explanations of Large Language Models Explain Language Representations in the Brain

SafetyDGX agent

arXiv:2502.14671v4 Announce Type: replace-cross Abstract: Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what driv

F^2Agent: Financial Fusion of Agentic Intelligence for Multimodal Trading

AgentsDGX agent

arXiv:2608.05668v1 Announce Type: cross Abstract: With increasingly diverse and heterogeneous information sources, effectively leveraging multimodal data is becoming pivotal for high-quality financial

Failing Gracefully: Mitigating Impact of Inevitable Robot Failures

SafetyDGX agent

arXiv:2608.05313v1 Announce Type: cross Abstract: Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failures such as

FI-TW: An Open Train-Weather Dataset for Railway Delay Analysis in Finland

SafetyDGX agent

arXiv:2601.16592v2 Announce Type: replace-cross Abstract: Train delays result from complex interactions between operational, technical, and environmental factors. While weather impacts railway reliabi

FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

Model ReleasesDGX agent

arXiv:2608.06144v1 Announce Type: new Abstract: Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution b

FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India

Model ReleasesDGX agent

arXiv:2608.06027v1 Announce Type: cross Abstract: In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write. Reaching them

From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction

ResearchDGX agent

arXiv:2608.05203v1 Announce Type: new Abstract: Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

SafetyDGX agent

arXiv:2608.06020v1 Announce Type: new Abstract: Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their belie

From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks

Local AiDGX agent

arXiv:2608.06227v1 Announce Type: cross Abstract: Despite advances in artificial intelligence (AI) across multiple sectors, today's AI tools, including deep learning and generative AI, still fail when

From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems

SafetyDGX agent

arXiv:2608.06112v1 Announce Type: new Abstract: Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc., yet most deployments remain isolated point solutions locked

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

Model ReleasesDGX agent

arXiv:2608.05948v1 Announce Type: new Abstract: Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit s

GAUGE: Granularity-Adaptive Counterfactual Gating of Evidence for Incomplete Multimodal Classification

ApplicationsDGX agent

arXiv:2608.05608v1 Announce Type: cross Abstract: Multimodal classification typically assumes all modalities are available, yet real-world inputs are often incomplete. Imputation and dynamic fusion ca

Gender-Based Heterogeneity in Youth Privacy-Protective Behavior for Smart Voice Assistants: Evidence from Multigroup PLS-SEM

ResearchDGX agent

arXiv:2603.27117v2 Announce Type: replace-cross Abstract: This paper investigates how gender shapes privacy decision-making in youth smart voice assistant (SVA) ecosystems. Using survey data from 469

GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.02721v3 Announce Type: replace Abstract: Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms the best

GROM: Gradient-Free Rapid One-Shot Machine Unlearning

Model ReleasesDGX agent

arXiv:2608.05783v1 Announce Type: cross Abstract: Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Current state

Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Baseline, and a Dual-Head Model

Model ReleasesDGX agent

arXiv:2608.05685v1 Announce Type: new Abstract: Most public benchmarks for machine-condition monitoring come from test rigs, where faults are induced on purpose and every event is known. Real producti

GSBF: Gaussian Splatting for Environment-Aware Beamforming

SafetyDGX agent

arXiv:2608.05896v1 Announce Type: new Abstract: Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design normally requires

Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture

Model ReleasesDGX agent

arXiv:2608.06130v1 Announce Type: cross Abstract: AI agents performing cryptographic operations (signing Git commits, authenticating API calls, issuing certificates) currently store private keys in so

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

Model ReleasesDGX agent

arXiv:2608.06301v1 Announce Type: new Abstract: As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts,

HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards

Model ReleasesDGX agent

arXiv:2608.06012v1 Announce Type: new Abstract: Search-agent rewards mix answer quality, citation grounding, tool cost, and anti-hacking terms; a high score therefore need not imply that cited evidenc

Hierarchical Latent Prediction for Language Models

ResearchDGX agent

arXiv:2608.05806v1 Announce Type: cross Abstract: While standard Next-Token Prediction (NTP) lays the foundation of language model pre- training, its teacher-forced training paradigm may not be optima

Hierarchical Server Architecture for Agentic Science

AgentsDGX agent

arXiv:2608.05332v1 Announce Type: cross Abstract: Agentic science is transforming the landscape of computational work, extending to scientific pipelines and workload managers. The workloads require sp

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots

Model ReleasesDGX agent

arXiv:2608.05715v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable

Hybrid Machine Learning Framework for Herd-Level Cattle Growth Pattern and Weight Gain Forecasting in Grazing-Based Production Systems

ApplicationsDGX agent

arXiv:2608.06001v1 Announce Type: new Abstract: Commercial grazing systems yield irregular livestock observations, which challenge cattle growth forecasting. This study developed a hybrid machine lear

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging

Model ReleasesDGX agent

arXiv:2608.05541v1 Announce Type: new Abstract: Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However,

HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection

ResearchDGX agent

arXiv:2608.05771v1 Announce Type: cross Abstract: Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance often degrades

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

AgentsDGX agent

arXiv:2608.06161v1 Announce Type: new Abstract: Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptu

IMMENSE: Inductive Multi-perspective User Classification in Social Networks

ApplicationsDGX agent

arXiv:2608.05259v1 Announce Type: cross Abstract: Online social networks increasingly expose people to users who propagate discriminatory, hateful, and violent content. Young users, in particular, are

Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks

SafetyDGX agent

arXiv:2608.05867v1 Announce Type: new Abstract: The use of ontologies and knowledge graphs is becoming increasingly widespread in the defence and national security domain. Numerous ontologies have bee

Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints

Model ReleasesDGX agent

arXiv:2608.06265v1 Announce Type: new Abstract: Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy

In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion

Local AiDGX agent

arXiv:2608.05237v1 Announce Type: cross Abstract: Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the curren

Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation

Model ReleasesDGX agent

arXiv:2608.05210v1 Announce Type: cross Abstract: Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by

Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability

Local AiDGX agent

arXiv:2608.05490v1 Announce Type: new Abstract: Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When s

Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis

TutorialsDGX agent

arXiv:2608.06037v1 Announce Type: new Abstract: Relational inductive biases are essential for capturing structural dependencies among data. This study investigates a dual-level relational framework fo

Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising

Model ReleasesDGX agent

arXiv:2510.05589v3 Announce Type: replace-cross Abstract: Effective time series forecasting enables various real-world applications, benefiting from the proliferation of mobile devices. However, the v

Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria

ApplicationsDGX agent

arXiv:2608.06364v1 Announce Type: cross Abstract: The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including fraud and reduced user control ove

Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

ResearchDGX agent

arXiv:2608.06122v1 Announce Type: cross Abstract: Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether simi

Large Language Models Threaten Double-blind Review

SafetyDGX agent

arXiv:2608.05157v1 Announce Type: cross Abstract: Double blind peer review serves as the scientific community primary defense against status and affiliation bias. Its effectiveness rests on the assump

Layer-wise Positional Bias in Short-Context Language Modeling

Model ReleasesDGX agent

arXiv:2601.04098v2 Announce Type: replace-cross Abstract: Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known as

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction

SafetyDGX agent

arXiv:2608.05600v1 Announce Type: cross Abstract: Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learn

Learning Context-Free Grammars for Grammar-Constrained Decoding via Declarative Agentic Programming with Guarantees

AgentsDGX agent

arXiv:2608.05493v1 Announce Type: cross Abstract: Language models (LMs) are increasingly used to interact with external services via programs written in domain-specific languages (DSLs). Unfortunately

Learning Globally Reusable Skills for Coding Agents

AgentsDGX agent

arXiv:2608.06153v1 Announce Type: cross Abstract: Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing approaches

Learning When to Trust via Selective Context Preference Optimization

Model ReleasesDGX agent

arXiv:2608.06377v1 Announce Type: cross Abstract: Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious rem

Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answering

Model ReleasesDGX agent

arXiv:2604.01280v2 Announce Type: replace-cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires Multimodal Large Language Models (MLLMs) to identify and combine fine-grained visu

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

Model ReleasesDGX agent

arXiv:2608.05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain beh

MACRO: Markov Chain Routing of Transformer Layers

SafetyDGX agent

arXiv:2608.05872v1 Announce Type: cross Abstract: Standard Large Language Models (LLMs) execute layers sequentially. Dynamic layer routing, i.e. search for a different execution path through layers in

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

Model ReleasesDGX agent

arXiv:2608.05850v1 Announce Type: cross Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, it

Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models

ApplicationsDGX agent

arXiv:2608.05243v1 Announce Type: cross Abstract: Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and inter

Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents

Model ReleasesDGX agent

arXiv:2606.21140v2 Announce Type: replace-cross Abstract: Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs d

Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers

Model ReleasesDGX agent

arXiv:2608.05472v1 Announce Type: cross Abstract: Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping

Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models

TutorialsDGX agent

arXiv:2608.05152v1 Announce Type: cross Abstract: Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior

Measuring and Detecting Harmful AI Sycophancy

TutorialsDGX agent

arXiv:2608.05624v1 Announce Type: new Abstract: Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This pa

← Previous
1…2324252627…354
Next →