AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
15 Jul 2026

Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability

Model ReleasesDGX agent

arXiv:2607.12056v1 Announce Type: new Abstract: Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and

Diversity-Enriched Option-Critic

TutorialsDGX agent

arXiv:2011.02565v2 Announce Type: replace-cross Abstract: Temporal abstraction allows reinforcement learning agents to represent knowledge and develop strategies over different temporal scales. The op

Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution

Model ReleasesDGX agent

arXiv:2607.13034v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task act


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory

Model ReleasesDGX agent

arXiv:2603.25112v2 Announce Type: replace-cross Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy)

Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?

Model ReleasesDGX agent

arXiv:2607.12787v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enab

Do We Really Need Transformers for Global Spatial Information Extraction in Traffic Forecasting?

ResearchDGX agent

arXiv:2607.12462v1 Announce Type: new Abstract: Existing traffic forecasting models commonly focus on extracting spatial dependencies, particularly global spatial information, which characterizes the

Do You Remember? Toward Memory-Centric Multimodal AI

ResearchDGX agent

arXiv:2607.11919v1 Announce Type: cross Abstract: Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability: they process images through a frozen v

Dynamic Resource Allocation for Ensemble Determinization MCTS

Model ReleasesDGX agent

arXiv:2607.13007v1 Announce Type: new Abstract: Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games with significant elements of randomn

Efficiently Learning Branching Networks for Multitask Algorithmic Reasoning

Model ReleasesDGX agent

arXiv:2512.01113v2 Announce Type: replace-cross Abstract: Algorithmic reasoning -- the ability to perform step-by-step logical inference -- is a synthetic benchmark for evaluating multi-step reasoning

Egocentric Bias in Vision-Language Models

Model ReleasesDGX agent

arXiv:2602.15892v2 Announce Type: replace-cross Abstract: Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet

Enabling 24-hour Agricultural Robotics: Unsupervised Day-to-Night Cross-Modal Image Translation for Nighttime Visual Navigation

Model ReleasesDGX agent

arXiv:2607.12065v1 Announce Type: cross Abstract: While visual navigation has been extensively studied in agricultural robotics, most existing systems assume daytime conditions. In fact, deploying aut

Enabling Energy-Efficient Simultaneous Multi-Task Reinforcement Learning through Spiking Neural Networks with Active Dendrites for Bio-inspired Generalist Agents

ApplicationsDGX agent

arXiv:2412.04847v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has demonstrated remarkable capabilities in training agents to solve complex tasks autonomously, such as mobile ro

Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)

Model ReleasesDGX agent

arXiv:2607.12336v1 Announce Type: cross Abstract: Artificial Intelligence (AI) technologies, while serving as a foundational enabler for modern social media and digital health services, exert a bivale

Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring

ResearchDGX agent

arXiv:2607.11981v1 Announce Type: cross Abstract: Aggregate reliability estimates can obscure heterogeneity in measurement-design burden across response conditions, so a single G- or D-study may misch

Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability

ApplicationsDGX agent

arXiv:2607.11963v1 Announce Type: cross Abstract: The early detection of Chronic Kidney Disease using machine learning has attracted significant interest in healthcare-related computer science. Despit

Evidence-Grounded AI for Musculoskeletal Care

Model ReleasesDGX agent

arXiv:2607.12527v1 Announce Type: new Abstract: Musculoskeletal diseases are among the leading causes of disability worldwide and create the greatest global need for rehabilitation. Because recovery,

Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

AgentsDGX agent

arXiv:2607.12650v1 Announce Type: cross Abstract: Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

ResearchDGX agent

arXiv:2509.22415v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains diffi

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Model ReleasesDGX agent

arXiv:2509.24372v3 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) has e

EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading

ResearchDGX agent

arXiv:2607.12455v1 Announce Type: new Abstract: Quantitative strategy optimization remains largely manual, requiring domain experts to identify weak signals, tune risk-control rules, and repeatedly va

Exact and Certified Data Shapley for Weighted k-Nearest-Neighbor Regression and Soft-Label Prediction

ResearchDGX agent

arXiv:2607.11956v1 Announce Type: cross Abstract: Data Shapley is the standard principled answer to which training points are worth what, and its k-nearest-neighbor (KNN) specialization is the version

Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction

Model ReleasesDGX agent

arXiv:2607.12584v1 Announce Type: cross Abstract: The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While rec

FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis

ApplicationsDGX agent

arXiv:2607.11464v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) addresses the limitations of Large Language Models (LLMs) when providing responses to domain-specific questions.

Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals

AgentsDGX agent

arXiv:2607.12233v1 Announce Type: cross Abstract: Large language model (LLM) trading agents show promising performance in equity markets, yet remain narrowly focused on US equities with little evidenc

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

Model ReleasesDGX agent

arXiv:2512.01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety

Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens

Local AiDGX agent

arXiv:2606.16847v3 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and quali

Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models

Local AiDGX agent

arXiv:2607.12962v1 Announce Type: cross Abstract: Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo controls in

FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation

Model ReleasesDGX agent

arXiv:2607.12982v1 Announce Type: new Abstract: Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic geometry remai

From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation

Model ReleasesDGX agent

arXiv:2607.12687v1 Announce Type: cross Abstract: LLMs can perform language-based quantitative prediction from unstructured inputs, but remain susceptible to hallucinations and overconfident errors, m

From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery

AgentsDGX agent

arXiv:2607.12474v1 Announce Type: new Abstract: Recent advances in foundation models have transformed AI for Science, enabling remarkably accurate predictive performance across domains ranging from pr

From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomography Data

ResearchDGX agent

arXiv:2607.12175v1 Announce Type: cross Abstract: X-ray tomography enables nondestructive characterization of material microstructures, while advances in micro-CT imaging have accelerated volumetric d

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

SafetyDGX agent

arXiv:2607.12463v1 Announce Type: new Abstract: Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in

G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis

SafetyDGX agent

arXiv:2607.11892v1 Announce Type: cross Abstract: Human-factor event diagnosis is essential for learning from operational events in nuclear power plants, yet its quality depends strongly on expert int

GaitSpan: Growing Humanoid Locomotion from Walking to Running

SafetyDGX agent

arXiv:2607.12114v1 Announce Type: cross Abstract: A humanoid that can walk should not relearn locomotion from scratch to jog or run. Yet current approaches often obtain gait diversity by prescribing g

Gene Expression-Informed Jointly Controlled Generative Modeling for Precision Molecular Design

ResearchDGX agent

arXiv:2607.11978v1 Announce Type: cross Abstract: Precision molecular design aims to discover personalized drug candidates through joint control of multiple conditions, such as biological relevance an

Generalized Distribution-Free Semi-Supervised Learning with Risk Rewrite

ResearchDGX agent

arXiv:2607.11947v1 Announce Type: cross Abstract: Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated. While PNU l

Git-Assistant: Planning-Based Support for Updating Git Repositories

SafetyDGX agent

arXiv:2607.09224v2 Announce Type: replace-cross Abstract: Version control systems are essential for collaborative software development, yet tools like git remain challenging for many practitioners. Re

Good Benchmarks

ResearchDGX agent

arXiv:2607.12217v1 Announce Type: new Abstract: Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons. The best tasks describe a real problem an experienced pr

Graph-Based Detection of Disinformation Narrative Diffusion between Russian and Ukrainian Telegram Channels

ResearchDGX agent

arXiv:2607.11894v1 Announce Type: cross Abstract: Detecting disinformation narratives on social media is challenging due to the scale of amplification, rapid evolution, and linguistic variability of o

Graph-Constrained Policy Learning for Extreme Clinical Code Prediction

Model ReleasesDGX agent

arXiv:2607.11954v1 Announce Type: cross Abstract: Clinical code prediction maps unstructured discharge summaries to ICD-10-CM leaf codes in a large, sparse, and deeply hierarchical label space. Most s

Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations

Local AiDGX agent

arXiv:2607.12077v1 Announce Type: new Abstract: Multi-agent language-model systems increasingly route local interactions, yet the runtime interaction graph is often treated as an implementation detail

GRID: Grammar-Railed Decoding for Enterprise SQL Generation

Model ReleasesDGX agent

arXiv:2607.11951v1 Announce Type: new Abstract: Large language models can write SQL, but enterprise deployment demands more than plausible text: outputs must be syntactically valid, must respect per-r

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

SafetyDGX agent

arXiv:2607.12752v1 Announce Type: cross Abstract: While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without expli

Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task

ResearchDGX agent

arXiv:2510.18315v2 Announce Type: replace-cross Abstract: We investigate how embedding dimension affects the emergence of an internal 'world model' in a transformer trained with reinforcement learning

How Inference Compute Shapes Frontier LLM Evaluation

Model ReleasesDGX agent

arXiv:2606.17930v2 Announce Type: replace Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result,

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

Model ReleasesDGX agent

arXiv:2607.12338v1 Announce Type: new Abstract: Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not sh

How Query Visibility Changes KV-Cache Compression Rankings: A Matched-Budget Audit

ResearchDGX agent

arXiv:2607.11942v1 Announce Type: cross Abstract: KV-cache compression methods are predominantly evaluated with the query appended to the context before compression -- a query-aware protocol. Yet the

HPC-Enabled Video-based Coastal Wave Parameter Estimation Using V-JEPA and Deep Spatiotemporal Learning

Model ReleasesDGX agent

arXiv:2607.11998v1 Announce Type: cross Abstract: High deployment cost, poor spatial coverage and susceptibility to storm conditions are all challenges faced by traditional in-situ methods. This paper

HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition

SafetyDGX agent

arXiv:2607.12774v1 Announce Type: cross Abstract: This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition. For multi-task learning with simultaneous p

Human-AI Agent Interaction as a Neuroplastic Training Environment

AgentsDGX agent

arXiv:2607.12823v1 Announce Type: new Abstract: Interaction with AI agents has become one of the most frequent activities of everyday digital life. Whether conversing with an assistant, working with a

IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation

AgentsDGX agent

arXiv:2607.10144v2 Announce Type: replace Abstract: Scientific ideation unfolds over multiple stages, including literature search, paper reading, tool use, claim checking, cross-paper synthesis, brain

I'm Sorry, but I Can't Help with Braille: Revealing Accessibility Failures in State-of-the-Art LLMs

SafetyDGX agent

arXiv:2607.11893v1 Announce Type: cross Abstract: Large Language Models (LLMs) perform strongly on many language tasks, but their capability in structurally constrained, accessibility-critical modalit

In-Context Reinforcement Learning under Non-Stationarity: A Survey

Model ReleasesDGX agent

arXiv:2607.11906v1 Announce Type: new Abstract: The development of decision-pretrained transformers, algorithm distillation, long-context meta-RL, and retrieval-augmented agents has renewed interest i

Inclusive Federated Learning Through Compliance-Weighted Noise Allocation in Healthcare AI

ApplicationsDGX agent

arXiv:2505.22108v4 Announce Type: replace-cross Abstract: Background: Federated learning (FL) enables collaborative training of clinical AI models without centralizing patient data, but adoption is li

Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration

SafetyDGX agent

arXiv:2607.12662v1 Announce Type: new Abstract: The paper introduces the Internet of Agentic Things (IoAT), an architectural framework that integrates agentic AI, IoT, cyber-physical systems, Physical

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement

AgentsDGX agent

arXiv:2606.19387v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved remarkable success in software development. However, they are susceptible to hallucinations, meanin

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

TutorialsDGX agent

arXiv:2607.11875v2 Announce Type: replace-cross Abstract: We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous wo

IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment

ResearchDGX agent

arXiv:2607.12375v1 Announce Type: cross Abstract: Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent approaches ba

Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

SafetyDGX agent

arXiv:2607.12406v1 Announce Type: new Abstract: The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequentl

JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks

SafetyDGX agent

arXiv:2602.06486v4 Announce Type: replace Abstract: Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, r

← Previous
1…7071727374…358
Next →