AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
20,981 results
11 Aug 2026

Application of Artificial Intelligence for Fraudulent Banking Operations Recognition

ResearchDGX agent

arXiv:2608.07471v1 Announce Type: cross Abstract: This study considers the task of applying artificial intelligence to recognize bank fraud. In recent years, due to the COVID19 pandemic, bank fraud ha

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups

Local AiDGX agent

arXiv:2608.08883v1 Announce Type: new Abstract: Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. H

ArchAgent v2: A Case Study with the Data Prefetching Championship

SafetyDGX agent

arXiv:2608.09874v1 Announce Type: new Abstract: Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarchitecture dis


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory

SafetyDGX agent

arXiv:2406.14373v3 Announce Type: replace Abstract: The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science

ASPaeroFlow: Decomposition Heuristics for Joint Air Traffic Flow & Capacity Management

Model ReleasesDGX agent

arXiv:2608.09315v1 Announce Type: new Abstract: While mathematical models act as vital decision support systems for operational Air Traffic Flow and Capacity Management (ATFCM), existing approaches is

ATLASFusion: Aggregation Tracking with Location-Aware Sparse Fusion for Robust Spatio-Temporal Multi-View Pedestrian Tracking

ResearchDGX agent

arXiv:2509.08421v2 Announce Type: replace-cross Abstract: For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintaining c

Attn-QAT: 4-Bit Attention With Quantization-Aware Training

ResearchDGX agent

arXiv:2603.00040v3 Announce Type: replace-cross Abstract: Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the ma

Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models

Model ReleasesDGX agent

arXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b

Automating Deception: Scalable Multi-Turn LLM Jailbreaks

Model ReleasesDGX agent

arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t

Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents

SafetyDGX agent

arXiv:2510.04465v3 Announce Type: replace-cross Abstract: LLM agents require personal information for personalization in order to effectively act on users' behalf, but this raises privacy concerns tha

AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts

Model ReleasesDGX agent

arXiv:2601.22758v2 Announce Type: replace Abstract: Large language model agents repeatedly encounter related tasks, yet systems that learn from trajectories commit every lesson to one predefined artif

Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks

Model ReleasesDGX agent

arXiv:2603.00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduc

Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics

Model ReleasesDGX agent

arXiv:2608.09638v1 Announce Type: new Abstract: Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reason

Back to the Future: A workbook time machine for spread sheet creation benchmarks

Model ReleasesDGX agent

arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived obj

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

Model ReleasesDGX agent

arXiv:2608.09888v1 Announce Type: cross Abstract: We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuou

Beyond 'I Can't Help With That': How Child Safety Experts Evaluate AI Chatbot Safety

SafetyDGX agent

arXiv:2608.07902v1 Announce Type: cross Abstract: Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes s

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Model ReleasesDGX agent

arXiv:2608.09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expecte

Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics

Model ReleasesDGX agent

arXiv:2512.05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily targe

Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents

Model ReleasesDGX agent

arXiv:2508.04412v3 Announce Type: replace Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

SafetyDGX agent

arXiv:2608.09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform tas

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

Model ReleasesDGX agent

arXiv:2603.12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge an

Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction

Model ReleasesDGX agent

arXiv:2608.08459v1 Announce Type: cross Abstract: Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domain

Beyond Uniform Restoration: Empowering All-in-One Restoration with Pixel-Level Multimodal Guidance

Local AiDGX agent

arXiv:2608.09482v1 Announce Type: cross Abstract: All-in-one image restoration is a unified low-level vision task that aims to effectively recover high-quality images from inputs degraded by various t

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

Model ReleasesDGX agent

arXiv:2608.09555v1 Announce Type: new Abstract: External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Yet their effe

Biologically Informed Representation Learning for Robust Cross-Center Generalization of MALDI-TOF Mass Spectrometry

Model ReleasesDGX agent

arXiv:2608.08182v1 Announce Type: cross Abstract: Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial identificati

Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models

AgentsDGX agent

arXiv:2512.11614v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuristic

BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference

ResearchDGX agent

arXiv:2608.07572v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their massive comput

Branch2Skill: Efficient Skill Evolution Through Reasoning Trees

AgentsDGX agent

arXiv:2608.08677v1 Announce Type: new Abstract: Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete o

Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search

Model ReleasesDGX agent

arXiv:2603.24084v2 Announce Type: replace Abstract: Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instances with i

Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

ApplicationsDGX agent

arXiv:2608.09045v1 Announce Type: cross Abstract: Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolate

Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts

Model ReleasesDGX agent

arXiv:2608.09510v1 Announce Type: cross Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and re

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

AgentsDGX agent

arXiv:2608.08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty,

Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline

AgentsDGX agent

arXiv:2608.09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questi

CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

Model ReleasesDGX agent

arXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, sup

Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization

ResearchDGX agent

arXiv:2608.08451v1 Announce Type: cross Abstract: Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reas

Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations

AgentsDGX agent

arXiv:2608.09268v1 Announce Type: cross Abstract: Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We s

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

Model ReleasesDGX agent

arXiv:2608.08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. Howev

Can Open-Weight Models Compete on Financial Text Comprehension?

Model ReleasesDGX agent

arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability

Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs

Model ReleasesDGX agent

arXiv:2608.08744v1 Announce Type: cross Abstract: The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds th

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

Model ReleasesDGX agent

arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven

Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

ResearchDGX agent

arXiv:2608.09485v1 Announce Type: new Abstract: Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission,

CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation

AgentsDGX agent

arXiv:2608.09790v1 Announce Type: new Abstract: Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions r

Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries

ApplicationsDGX agent

arXiv:2608.09532v1 Announce Type: cross Abstract: Enterprises increasingly seek to query data lakes using natural language via AI-driven tools like semantic operators or deep research agents. However,

CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift

Model ReleasesDGX agent

arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both ha

CDGC-Net: 3D Medical Image Segmentation with Cooperative Dual-Scale Self-Attention and Grouped Channel Modeling

SafetyDGX agent

arXiv:2608.08575v1 Announce Type: cross Abstract: Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing methods often mod

CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems

AgentsDGX agent

arXiv:2608.09848v1 Announce Type: new Abstract: The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a c

CFD-Guided Detection of Concept Drift in Multimodal Physiologic Signals

ResearchDGX agent

arXiv:2608.07759v1 Announce Type: cross Abstract: Cardiovascular AI models can classify clean elec- trocardiogram (ECG) signals, but real wearable signals change because of motion, breathing, posture,

ChronoState: Hidden Elapsed-Time Conditioning for Temporal-State Action Selection in Frozen-Backbone Language Models

Model ReleasesDGX agent

arXiv:2608.09124v1 Announce Type: new Abstract: Temporal decisions in language-model systems often depend on both symbolic task state and elapsed wall-clock time, such as cache expiration, job complet

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

Model ReleasesDGX agent

arXiv:2608.09164v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms.

CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

Model ReleasesDGX agent

arXiv:2608.09374v1 Announce Type: new Abstract: Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, sel

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

Model ReleasesDGX agent

arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investiga

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primari

Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth

SafetyDGX agent

arXiv:2608.07564v1 Announce Type: cross Abstract: In digital dentistry and oral surgery, the registration of jawbone CT and intraoral scanner (IOS) data is essential for integrating internal bone stru

ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners

Local AiDGX agent

arXiv:2608.09732v1 Announce Type: cross Abstract: Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find th

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

Model ReleasesDGX agent

arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise

COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping

Model ReleasesDGX agent

arXiv:2608.07570v1 Announce Type: cross Abstract: Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-

Communication-efficient distributed hazard difference estimation for heterogeneous multi-site survival data

ResearchDGX agent

arXiv:2601.14609v2 Announce Type: replace-cross Abstract: Multi-site collaboration can power survival models that no single hospital could fit alone, but privacy rules and protected computing environm

Complete, Scalable, and Robust Prioritized Planning for Multi-Robot Ordered Storage and Retrieval at Maximum Capacity

AgentsDGX agent

arXiv:2608.07734v1 Announce Type: cross Abstract: Automated warehouses face a fundamental trade-off between maximizing storage density and achieving high retrieval throughput. While puzzle-based stora

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

TutorialsDGX agent

arXiv:2608.08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitatio

Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario

AgentsDGX agent

arXiv:2608.08131v1 Announce Type: cross Abstract: In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short directive activat

← Previous
1…45678…350
Next →