AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
28 Apr 2026

Benchmarking Emergent Coordination in Large-Scale LLM Populations: An Evaluation Framework on the MoltBook Archive

Model ReleasesDGX agent

arXiv:2603.03555v2 Announce Type: replace-cross Abstract: As multi-agent Large Language Model (LLM) systems scale, evaluating their emergent coordination dynamics becomes increasingly critical. Howeve

Benchmarking Source-Sensitive Reasoning in Turkish: Humans and LLMs under Evidential Trust Manipulation

Local AiDGX agent

arXiv:2604.24665v1 Announce Type: cross Abstract: This paper investigates whether source trustworthiness shapes Turkish evidential morphology and whether large language models (LLMs) track this sensit

BERT-APC: A Reference-free Framework for Automatic Pitch Correction via Musical Context Inference

Research

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2511.20006v2 Announce Type: replace-cross Abstract: Automatic Pitch Correction (APC) enhances vocal recordings by aligning pitch deviations with intended musical notes. However, existing APC sys

Beyond Context: Large Language Models' Failure to Grasp Users' Intent

Model ReleasesDGX agent

arXiv:2512.21110v3 Announce Type: replace Abstract: Current Large Language Models (LLMs) safety approaches focus on explicitly harmful content while overlooking a critical vulnerability: the inability

Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models

SafetyDGX agent

arXiv:2502.14888v4 Announce Type: replace-cross Abstract: The success of vision-language models is primarily attributed to effective alignment across modalities such as vision and language. However, m

Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs

TutorialsDGX agent

arXiv:2602.02556v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are largely static and often redo reasoning or repeat mistakes. Prior experience reuse typically relies on extern

Beyond Single-Agent Alignment: Preventing Context-Fragmented Violations in Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2604.22879v1 Announce Type: cross Abstract: We identify and formalize a novel security risk: Context-Fragmented Violations (CFVs) - a class of policy breaches where individual agent actions appe

Beyond Static: Related Questions Retrieval Through Conversations in Community Question Answering

TutorialsDGX agent

arXiv:2604.22759v1 Announce Type: cross Abstract: In community question answering (cQA) platforms like Stack Overflow, related question retrieval is recognized as a fundamental task that allows users

Beyond the Attention Stability Boundary: Agentic Self-Synthesizing Reasoning Protocols

Model ReleasesDGX agent

arXiv:2604.24512v1 Announce Type: new Abstract: As LLM agents transition to autonomous digital coworkers, maintaining deterministic goal-directedness in non-linear multi-turn conversations emerged as

BiTA: Bidirectional Gated Recurrent Unit-Transformer Aggregator in a Temporal Graph Network Framework for Alert Prediction in Computer Networks

ApplicationsDGX agent

arXiv:2604.22781v1 Announce Type: cross Abstract: Proactive alert prediction in computer networks is critical for mitigating evolving cyber threats and enabling timely defensive actions. Temporal Grap

BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks

AgentsDGX agent

arXiv:2508.08127v2 Announce Type: replace Abstract: The security of LLM-based multi-agent systems (MAS) is critically threatened by propagation vulnerability, where malicious agents can distort collec

C-MORAL: Controllable Multi-Objective Molecular Optimization with Reinforcement Alignment for LLMs

Model ReleasesDGX agent

arXiv:2604.23061v1 Announce Type: cross Abstract: Large language models (LLMs) show promise for molecular optimization, but aligning them with selective and competing drug-design constraints remains c

Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft

Model ReleasesDGX agent

arXiv:2604.24697v1 Announce Type: new Abstract: Discovering causal regularities and applying them to build functional systems--the discovery-to-application loop--is a hallmark of general intelligence,

Can Large Language Models Really Recognize Your Name?

Model ReleasesDGX agent

arXiv:2505.14549v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly being used in privacy pipelines to detect and remedy sensitive data leakage. These solutions oft

Can Multimodal Large Language Models Truly Understand Small Objects?

Model ReleasesDGX agent

arXiv:2604.22884v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown promising potential in diverse understanding tasks, e.g., image and video analysis, math and physi

CAP-CoT: Cycle Adversarial Prompt for Improving Chain of Thoughts in LLM Reasoning

SafetyDGX agent

arXiv:2604.23270v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has emerged as a simple and effective way to elicit step-by-step solutions from large language models (LLMs). However,

CAPSULE: Control-Theoretic Action Perturbations for Safe Uncertainty-Aware Reinforcement Learning

SafetyDGX agent

arXiv:2604.23576v1 Announce Type: cross Abstract: Ensuring safe exploration in high-dimensional systems with unknown dynamics remains a significant challenge. Existing safe reinforcement learning meth

CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation

ResearchDGX agent

arXiv:2601.06352v2 Announce Type: replace Abstract: Adapting large language models to individual users remains challenging due to the tension between fine-grained personalization and scalable deployme

Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters

AgentsDGX agent

arXiv:2604.24710v1 Announce Type: new Abstract: Objective. Clinical AI documentation systems require evaluation methodologies that are clinically valid, economically viable, and sensitive to iterative

Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis

Model ReleasesDGX agent

arXiv:2510.16371v2 Announce Type: replace-cross Abstract: The development of computer-assisted surgery systems relies on large-scale, annotated datasets. Existing cataract surgery resources lack the d

Causal Discovery as Dialectical Aggregation: A Quantitative Argumentation Framework

Model ReleasesDGX agent

arXiv:2604.23633v1 Announce Type: new Abstract: Constraint-based causal discovery is brittle in finite-sample regimes because erroneous conditional-independence (CI) decisions can cascade into substan

Certified geometric robustness -- Super-DeepG

SafetyDGX agent

arXiv:2604.24379v1 Announce Type: new Abstract: Safety-critical applications are required to perform as expected in normal operations. Image processing functions are often required to be insensitive t

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies

Local AiDGX agent

arXiv:2604.24622v1 Announce Type: cross Abstract: Flow-based vision-language-action (VLA) policies offer strong expressivity for action generation, but suffer from a fundamental inefficiency: multi-st

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics

Model ReleasesDGX agent

arXiv:2509.20374v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated strong performance across general NLP tasks, but their utility in automating numerical experime

Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment

ResearchDGX agent

arXiv:2604.24447v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are promising for generalist robot control, but on-robot deployment is bottlenecked by real-time inference under t

CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging

ResearchDGX agent

arXiv:2604.22989v1 Announce Type: cross Abstract: Recent medical multimodal foundation models are built as multimodal LLMs (MLLMs) by connecting a CLIP-pretrained vision encoder to an LLM using LLaVA-

Citation-Driven Multi-View Training for Patent Embeddings: QaECTER and Sophia-Bench

Model ReleasesDGX agent

arXiv:2604.22897v1 Announce Type: cross Abstract: Patent retrieval underpins critical decisions in innovation, examination, and IP strategy, yet progress has been hampered by the absence of benchmarks

ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation

Model ReleasesDGX agent

arXiv:2604.23853v1 Announce Type: new Abstract: Skill-distillation pipelines learn reusable rules from LLM agent trajectories, but they lack a key signal: how much each step costs. Without per-step co

ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis

Model ReleasesDGX agent

arXiv:2604.16922v2 Announce Type: replace Abstract: Climate research is pivotal for mitigating global environmental crises, yet the accelerating volume of multi-scale datasets and the complexity of an

CLIN-LLM: A Safety-Constrained Hybrid Framework for Clinical Diagnosis and Treatment Generation

Model ReleasesDGX agent

arXiv:2510.22609v2 Announce Type: replace Abstract: Accurate symptom-to-disease classification and clinically grounded treatment recommendations remain challenging, particularly in heterogeneous patie

Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training

ApplicationsDGX agent

arXiv:2303.05330v1 Announce Type: cross Abstract: Geo-distributed ML training can benefit many emerging ML scenarios (e.g., large model training, federated learning) with multi-regional cloud resource

CNN-ViT Fusion with Adaptive Attention Gate for Brain Tumor MRI Classification: A Hybrid Deep Learning Model

Local AiDGX agent

arXiv:2604.23137v1 Announce Type: cross Abstract: Early detection and classifying brain tumors using Magnetic Resonance Imaging (MRI) images is highly important but difficult to extract in medical ima

Code Broker: A Multi-Agent System for Automated Code Quality Assessment

Local AiDGX agent

arXiv:2604.23088v1 Announce Type: cross Abstract: We present Code Broker, a multi agent system built with Google Agent Development Kit ADK that analyses Python code from files, local directories, or G

CombiMOTS: Combinatorial Multi-Objective Tree Search for Dual-Target Molecule Generation

SafetyDGX agent

arXiv:2604.23307v1 Announce Type: cross Abstract: Dual-target molecule generation, which focuses on discovering compounds capable of interacting with two target proteins, has garnered significant atte

COMO: Closed-Loop Optical Molecule Recognition with Minimum Risk Training

SafetyDGX agent

arXiv:2604.23546v1 Announce Type: cross Abstract: Optical chemical structure recognition (OCSR) translates molecular images into machine-readable representations like SMILES strings or molecular graph

Comparative Insights on Adversarial Machine Learning from Industry and Academia: A User-Study Approach

ApplicationsDGX agent

arXiv:2602.04753v2 Announce Type: replace-cross Abstract: An exponential growth of Machine Learning and its Generative AI applications brings with it significant security challenges, often referred to

Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent Workflows

Model ReleasesDGX agent

arXiv:2604.22820v1 Announce Type: cross Abstract: Long-horizon tool-using tasks sometimes benefit from revisiting earlier subtasks for recovery and exploration, but added multi-agent workflow flexibil

Conformal PM2.5 Mapping Under Spatial Covariate Shift: Satellite-Reanalysis Fusion for Africa's Green Industrial Transition

ResearchDGX agent

arXiv:2604.22787v1 Announce Type: cross Abstract: Africa's green industrialization imperative demands reliable infrastructure for monitoring air quality. We present a satellite-reanalysis PM2.5 fusion

ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation

SafetyDGX agent

arXiv:2504.02316v4 Announce Type: replace-cross Abstract: Recent advances in zero-shot text-to-3D generation have revolutionized 3D content creation by enabling direct synthesis from textual descripti

Constraint-Based Analysis of Reasoning Shortcuts in Neurosymbolic Learning

Model ReleasesDGX agent

arXiv:2604.23377v1 Announce Type: new Abstract: Neurosymbolic systems can satisfy logical constraints during learning without achieving the intended concept-label correspondence; this is a problem kno

Constraint-Guided Multi-Agent Decompilation for Executable Binary Recovery

AgentsDGX agent

arXiv:2604.23940v1 Announce Type: cross Abstract: Decompilation -- recovering source code from compiled binaries -- is essential for security analysis, malware reverse engineering, and legacy software

Context-Aware Hospitalization Forecasting Evaluations for Decision Support using LLMs

SafetyDGX agent

arXiv:2604.23949v1 Announce Type: new Abstract: Medical and public health experts must make real-time resource decisions, such as expanding hospital bed capacity, based on projected hospitalization tr

Cooperative Informative Sensing for Monitoring Dynamic Indoor Environments via Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2604.23179v1 Announce Type: cross Abstract: Monitoring human activity in indoor environments is important for applications such as facility management, safety assessment, and space utilization a

CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge Deployment

ResearchDGX agent

arXiv:2410.13903v3 Announce Type: replace-cross Abstract: Proprietary large language models (LLMs) exhibit strong generalization capabilities across diverse tasks and are increasingly deployed on edge

CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning

Model ReleasesDGX agent

arXiv:2601.14952v2 Announce Type: replace-cross Abstract: While large language models now handle million-token contexts, their capacity for reasoning across entire document repositories remains largel

Cortex-Inspired Continual Learning: Unsupervised Instantiation and Recovery of Functional Task Networks

Model ReleasesDGX agent

arXiv:2604.24637v1 Announce Type: cross Abstract: Block-sequential continual learning demands that a single model both protect prior solutions from catastrophic forgetting and efficiently infer at inf

Credal Concept Bottleneck Models for Epistemic-Aleatoric Uncertainty Decomposition

ResearchDGX agent

arXiv:2604.24170v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) predict through human-interpretable concepts, but they typically output point concept probabilities that conflate epist

CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks

Model ReleasesDGX agent

arXiv:2510.17687v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) achieve strong reasoning and perception capabilities but are increasingly vulnerable to jailbreak att

Crystal structure prediction using graph neural combinatorial optimization

HardwareDGX agent

arXiv:2604.23921v1 Announce Type: cross Abstract: Crystalline materials are widely used in technological applications, yet their discovery remains a significant challenge. As their properties are driv

CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation

Model ReleasesDGX agent

arXiv:2604.24001v1 Announce Type: new Abstract: The evaluation of generated reports remains a critical challenge in Computed Tomography (CT) report generation, due to the large volume of text, the div

CT-Guided Spatially-varying Regularization for Voxel-Wise Deformable Whole-Body PET Registration

SafetyDGX agent

arXiv:2604.22905v1 Announce Type: cross Abstract: Whole-body Positron Emission Tomography (PET) registration is essential for multi-parametric tumor characterization and assessment of metastatic disea

CUB: Benchmarking Context Utilisation Techniques for Language Models

Model ReleasesDGX agent

arXiv:2505.16518v3 Announce Type: replace-cross Abstract: Incorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language mod

CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning

SafetyDGX agent

arXiv:2601.13262v2 Announce Type: replace Abstract: While large language models (LLMs) have shown to perform well on monolingual mathematical and commonsense reasoning, they remain unreliable for mult

CyberCane: Neuro-Symbolic RAG for Privacy-Preserving Phishing Detection with Formal Ontology Reasoning

ApplicationsDGX agent

arXiv:2604.23563v1 Announce Type: cross Abstract: Privacy-critical domains require phishing detection systems that satisfy contradictory constraints: near-zero false positives to prevent workflow disr

Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech

SafetyDGX agent

arXiv:2510.05799v2 Announce Type: replace-cross Abstract: Aligning text-to-speech (TTS) system outputs with human feedback through preference optimization has been shown to effectively improve the rob

Decoding the mechanisms of the Hattrick football manager game using Bayesian network structure learning

ResearchDGX agent

arXiv:2504.09499v2 Announce Type: replace-cross Abstract: Hattrick is a free web-based probabilistic football manager game with over 200,000 users competing for titles at national and international le

DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting

Model ReleasesDGX agent

arXiv:2604.23968v1 Announce Type: cross Abstract: Accurate time series forecasting in scientific domains such as climate modeling, physiological monitoring, and energy systems benefits from both compe

Deep Learning-Enabled Dissolved Oxygen Sensing in Biofouling Environments for Ocean Monitoring

Model ReleasesDGX agent

arXiv:2604.24236v1 Announce Type: cross Abstract: The escalating climate crisis and ecosystem degradation demand intelligent, low-cost sensors capable of robust, long-term monitoring in real-world env

DeepImagine: Learning Biomedical Reasoning via Successive Counterfactual Imagining

Model ReleasesDGX agent

arXiv:2604.23054v1 Announce Type: cross Abstract: Predicting the outcomes of prospective clinical trials remains a major challenge for large language models. Prior work has shown that both traditional

DeepSignature: Digitally Signed, Content-Encoding Watermarks for Robust and Transparent Image Authentication

Local AiDGX agent

arXiv:2604.23016v1 Announce Type: cross Abstract: AI-powered generative models have significantly expanded the possibilities for editing, manipulating, and creating high-quality images. Particularly,

← Previous
1…296297298299300…354
Next →