AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
24 Apr 2026

Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models

AgentsDGX agent

arXiv:2604.20994v1 Announce Type: cross Abstract: The growth of agentic AI has drawn significant attention to function calling Large Language Models (LLMs), which are designed to extend the capabiliti

Bridging the Training-Deployment Gap: Gated Encoding and Multi-Scale Refinement for Efficient Quantization-Aware Image Enhancement

ResearchDGX agent

arXiv:2604.21743v1 Announce Type: new Abstract: Image enhancement models for mobile devices often struggle to balance high output quality with the fast processing speeds required by mobile hardware. W

Brief chatbot interactions produce lasting changes in human moral values

AgentsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.21430v1 Announce Type: new Abstract: Moral judgements form the foundation of human social behavior and societal systems. While Artificial Intelligence chatbots increasingly serve as persona

Building a Precise Video Language with Human-AI Oversight

Model ReleasesDGX agent

arXiv:2604.21718v1 Announce Type: cross Abstract: Video-language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmark

C-SHAP for time series: An approach to high-level temporal explanations

ApplicationsDGX agent

arXiv:2504.11159v2 Announce Type: replace Abstract: In high-stakes domains, such as healthcare and industry, the explainability of AI-based decision-making has become crucial. Without insight into mod

Calibeating Prediction-Powered Inference

ResearchDGX agent

arXiv:2604.21260v1 Announce Type: cross Abstract: We study semisupervised mean estimation with a small labeled sample, a large unlabeled sample, and a black-box prediction model whose output may be mi

Can MLLMs 'Read' What is Missing?

Model ReleasesDGX agent

arXiv:2604.21277v1 Announce Type: new Abstract: We introduce MMTR-Bench, a benchmark designed to evaluate the intrinsic ability of Multimodal Large Language Models (MLLMs) to reconstruct masked text d

CAP: Controllable Alignment Prompting for Unlearning in LLMs

Model ReleasesDGX agent

arXiv:2604.21251v1 Announce Type: cross Abstract: Large language models (LLMs) trained on unfiltered corpora inherently risk retaining sensitive information, necessitating selective knowledge unlearni

CaST-POI: Candidate-Conditioned Spatiotemporal Modeling for Next POI Recommendation

Model ReleasesDGX agent

arXiv:2604.20845v1 Announce Type: cross Abstract: Next Point-of-Interest (POI) recommendation plays a crucial role in location-based services by predicting users' future mobility patterns. Existing me

Causal Disentanglement for Full-Reference Image Quality Assessment

ResearchDGX agent

arXiv:2604.21654v1 Announce Type: cross Abstract: Existing deep network-based full-reference image quality assessment (FR-IQA) models typically work by performing pairwise comparisons of deep features

ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models

ResearchDGX agent

arXiv:2509.24239v4 Announce Type: replace-cross Abstract: Recent large language models (LLMs) have shown strong reasoning capabilities. However, a critical question remains: do these models possess ge

Clinical Reasoning AI for Oncology Treatment Planning: A Multi-Specialty Case-Based Evaluation

SafetyDGX agent

arXiv:2604.20869v1 Announce Type: cross Abstract: Background: More than 80% of U.S. cancer care is delivered in community settings, where survival remains worse than at academic centers. Clinicians mu

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks

AgentsDGX agent

arXiv:2604.20987v1 Announce Type: new Abstract: Long horizon interactive environments are a testbed for evaluating agents skill usage abilities. These environments demand multi step reasoning, the cha

CoFEE: Reasoning Control for LLM-Based Feature Discovery

ResearchDGX agent

arXiv:2604.21584v1 Announce Type: new Abstract: Feature discovery from complex unstructured data is fundamentally a reasoning problem: it requires identifying abstractions that are predictive of a tar

Cognitive Amplification vs Cognitive Delegation in Human-AI Systems: A Metric Framework

Model ReleasesDGX agent

arXiv:2603.18677v2 Announce Type: replace-cross Abstract: Artificial intelligence is increasingly embedded in human decision making. In some cases, it enhances human reasoning. In others, it fosters e

Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning

SafetyDGX agent

arXiv:2509.23744v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) promise enhanced reasoning by integrating diverse inputs such as text, vision, and audio. Yet cross-m

Conjecture and Inquiry: Quantifying Software Performance Requirements via Interactive Retrieval-Augmented Preference Elicitation

ApplicationsDGX agent

arXiv:2604.21380v1 Announce Type: cross Abstract: Since software performance requirements are documented in natural language, quantifying them into mathematical forms is essential for software enginee

Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs

Model ReleasesDGX agent

arXiv:2509.21361v2 Announce Type: replace-cross Abstract: Large language model (LLM) providers boast big numbers for maximum context window sizes. To test the real world use of context windows, we 1)

Continuous-Utility Direct Preference Optimization

SafetyDGX agent

arXiv:2602.00931v2 Announce Type: replace-cross Abstract: Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture par

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors

Model ReleasesDGX agent

arXiv:2604.21241v1 Announce Type: cross Abstract: Vision--Language--Action (VLA) models often use intermediate representations to connect multimodal inputs with continuous control, yet spatial guidanc

Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination

Model ReleasesDGX agent

arXiv:2506.21546v4 Announce Type: replace-cross Abstract: Segmentation Vision-Language Models (VLMs) have significantly advanced grounded visual understanding, yet they remain prone to pixel-grounding

Cross-Entropy Is Load-Bearing: A Pre-Registered Scope Test of the K-Way Energy Probe on Bidirectional Predictive Coding

Model ReleasesDGX agent

arXiv:2604.21286v1 Announce Type: cross Abstract: Cacioli (2026) showed that the K-way energy probe on standard discriminative predictive coding networks reduces approximately to a monotone function o

Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms

Model ReleasesDGX agent

arXiv:2604.21131v1 Announce Type: cross Abstract: AI-agent guardrails are memoryless: each message is judged in isolation, so an adversary who spreads a single attack across dozens of sessions slips p

Crystal: Characterizing Relative Impact of Scholarly Publications

SafetyDGX agent

arXiv:2603.26791v2 Announce Type: replace-cross Abstract: Assessing a cited paper's impact is typically done by analyzing its citation context in isolation within the citing paper. While this focuses

CSC: Turning the Adversary's Poison against Itself

Model ReleasesDGX agent

arXiv:2604.21416v1 Announce Type: cross Abstract: Poisoning-based backdoor attacks pose significant threats to deep neural networks by embedding triggers in training data, causing models to misclassif

Data-Driven Open-Loop Simulation for Digital-Twin Operator Decision Support in Wastewater Treatment

Model ReleasesDGX agent

arXiv:2604.20935v1 Announce Type: cross Abstract: Wastewater treatment plants (WWTPs) need digital-twin-style decision support tools that can simulate plant response under prescribed control plans, to

Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research

Model ReleasesDGX agent

arXiv:2604.21006v1 Announce Type: new Abstract: We introduce Deep FinResearch Bench, a practical and comprehensive evaluation framework for deep research (DR) agents in financial investment research.

Deep Interest Mining with Cross-Modal Alignment for SemanticID Generation in Generative Recommendation

SafetyDGX agent

arXiv:2604.20861v1 Announce Type: cross Abstract: Generative Recommendation (GR) has demonstrated remarkable performance in next-token prediction paradigms, which relies on Semantic IDs (SIDs) to comp

DenoiseRank: Learning to Rank by Diffusion Models

Model ReleasesDGX agent

arXiv:2604.20852v1 Announce Type: cross Abstract: Learning to rank (LTR) is one of the core tasks in Machine Learning. Traditional LTR models have made great progress, but nearly all of them are imple

DiagramBank: A Large-scale Dataset of Diagram Design Exemplars with Paper Metadata for Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2604.20857v1 Announce Type: cross Abstract: Recent advances in autonomous ``AI scientist'' systems have demonstrated the ability to automatically write scientific manuscripts and codes with exec

Dialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles

Model ReleasesDGX agent

arXiv:2604.21152v1 Announce Type: cross Abstract: As state-of-the-art Large Language Models (LLMs) have become ubiquitous, ensuring equitable performance across diverse demographics is critical. Howev

Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation

ApplicationsDGX agent

arXiv:2604.21421v1 Announce Type: cross Abstract: Protecting patient privacy in clinical narratives is essential for enabling secondary use of healthcare data under regulations such as GDPR and HIPAA.

Differentially Private Model Merging

Model ReleasesDGX agent

arXiv:2604.20985v1 Announce Type: cross Abstract: In machine learning applications, privacy requirements during inference or deployment time could change constantly due to varying policies, regulation

Dilated CNNs for Periodic Signal Processing: A Low-Complexity Approach

ResearchDGX agent

arXiv:2604.21651v1 Announce Type: cross Abstract: Denoising of periodic signals and accurate waveform estimation are core tasks across many signal processing domains, including speech, music, medical

Divide-then-Diagnose: Weaving Clinician-Inspired Contexts for Ultra-Long Capsule Endoscopy Videos

ResearchDGX agent

arXiv:2604.21814v1 Announce Type: cross Abstract: Capsule endoscopy (CE) enables non-invasive gastrointestinal screening, but current CE research remains largely limited to frame-level classification

Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition

SafetyDGX agent

arXiv:2604.21276v1 Announce Type: cross Abstract: As pretrained large language models replace task-specific decoders in speech recognition, a critical question arises: do their text-derived priors mak

Domain-Aware Hierarchical Contrastive Learning for Semi-Supervised Generalization Fault Diagnosis

Model ReleasesDGX agent

arXiv:2604.20928v1 Announce Type: cross Abstract: Fault diagnosis under unseen operating conditions remains highly challenging when labeled data are scarce. Semi-supervised domain generalization fault

Doubly Saturated Ramsey Graphs: A Case Study in Computer-Assisted Mathematical Discovery

ApplicationsDGX agent

arXiv:2604.21187v1 Announce Type: cross Abstract: Ramsey-good graphs are graphs that contain neither a clique of size s nor an independent set of size t. We study doubly saturated Ramsey-good graphs,

Drug Synergy Prediction via Residual Graph Isomorphism Networks and Attention Mechanisms

Model ReleasesDGX agent

arXiv:2604.21473v1 Announce Type: cross Abstract: In the treatment of complex diseases, treatment regimens using a single drug often yield limited efficacy and can lead to drug resistance. In contrast

DryRUN: On the Role of Public Tests in LLM-Driven Code Generation

AgentsDGX agent

arXiv:2604.21598v1 Announce Type: cross Abstract: Multi-agent frameworks are widely used in autonomous code generation and have applications in complex algorithmic problem-solving. Recent work has add

Dynamical Priors as a Training Objective in Reinforcement Learning

SafetyDGX agent

arXiv:2604.21464v1 Announce Type: cross Abstract: Standard reinforcement learning (RL) optimizes policies for reward but imposes few constraints on how decisions evolve over time. As a result, policie

Efficient Agent Evaluation via Diversity-Guided User Simulation

AgentsDGX agent

arXiv:2604.21480v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as customer-facing agents, yet evaluating their reliability remains challenging due to stochastic

Efficient Logic Gate Networks for Video Copy Detection

ResearchDGX agent

arXiv:2604.21694v1 Announce Type: cross Abstract: Video copy detection requires robust similarity estimation under diverse visual distortions while operating at very large scale. Although deep neural

Empirical Comparison of Agent Communication Protocols for Task Orchestration

Model ReleasesDGX agent

arXiv:2603.22823v3 Announce Type: replace Abstract: Context. The problem of comparative evaluation of communication protocols for task orchestration by large language model (LLM) agents is considered.

Enabling and Inhibitory Pathways of University Students' Willingness to Disclose AI Use: A Cognition-Affect-Conation Perspective

SafetyDGX agent

arXiv:2604.21733v1 Announce Type: new Abstract: The increasing integration of artificial intelligence (AI) in higher education has raised important questions regarding students' transparency in report

Engaged AI Governance: Addressing the Last Mile Challenge Through Internal Expert Collaboration

SafetyDGX agent

arXiv:2604.21554v1 Announce Type: new Abstract: Under the EU AI Act, translating AI governance requirements into software development practice remains challenging. While AI governance frameworks exist

EngramaBench: Evaluating Long-Term Conversational Memory with Structured Graph Retrieval

Model ReleasesDGX agent

arXiv:2604.21229v1 Announce Type: cross Abstract: Large language model assistants are increasingly expected to retain and reason over information accumulated across many sessions. We introduce Engrama

Enhancing Online Recruitment with Category-Aware MoE and LLM-based Data Augmentation

ResearchDGX agent

arXiv:2604.21264v1 Announce Type: new Abstract: Person-Job Fit (PJF) is a critical component for online recruitment. Existing approaches face several challenges, particularly in handling low-quality j

Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reasoning-Component Classification

Model ReleasesDGX agent

arXiv:2604.21137v1 Announce Type: cross Abstract: Analyzing the reasoning patterns of students in science classrooms is critical for understanding knowledge construction mechanism and improving instru

Equity Bias: An Ethical Framework for AI Design

SafetyDGX agent

arXiv:2604.21907v1 Announce Type: cross Abstract: Equity Bias is a philosophical and practical framework for building smarter, more equitable AI systems. Grounded in hermeneutic philosophy and epistem

ERA: Evidence-based Reliability Alignment for Honest Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2604.20854v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) grounds language models in factual evidence but introduces critical challenges regarding knowledge conflicts betw

Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI

SafetyDGX agent

arXiv:2604.20972v1 Announce Type: new Abstract: Content moderation systems are typically evaluated by measuring agreement with human labels. In rule-governed environments this assumption fails: multip

Evaluating AI Meeting Summaries with a Reusable Cross-Domain Pipeline

Model ReleasesDGX agent

arXiv:2604.21345v1 Announce Type: new Abstract: We present a reusable evaluation pipeline for generative AI applications, instantiated for AI meeting summaries and released with a public artifact pack

Expanding the extreme-k dielectric materials space through physics-validated generative reasoning

ResearchDGX agent

arXiv:2604.21068v1 Announce Type: cross Abstract: The most technologically consequential materials are often the rarest: they occupy narrow regions of chemical space, obey competing physical constrain

Exploring the Role of Synthetic Data Augmentation in Controllable Human-Centric Video Generation

ApplicationsDGX agent

arXiv:2604.21291v1 Announce Type: cross Abstract: Controllable human video generation aims to produce realistic videos of humans with explicitly guided motions and appearances,serving as a foundation

Fairness Evaluation and Inference Level Mitigation in LLMs

SafetyDGX agent

arXiv:2510.18914v4 Announce Type: replace-cross Abstract: Large language models often display undesirable behaviors embedded in their internal representations, undermining fairness, inconsistency drif

Fairness under uncertainty in sequential decisions

SafetyDGX agent

arXiv:2604.21711v1 Announce Type: cross Abstract: Fair machine learning (ML) methods help identify and mitigate the risk that algorithms encode or automate social injustices. Algorithmic approaches al

FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation

SafetyDGX agent

arXiv:2604.21420v1 Announce Type: new Abstract: Quality Estimation (QE) aims to assess machine translation quality without reference translations, but recent studies have shown that existing QE models

Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding

Model ReleasesDGX agent

arXiv:2506.19579v3 Announce Type: replace-cross Abstract: Robotic scene understanding increasingly relies on Vision-Language Models (VLMs) to generate natural language descriptions of the environment.

Federated Co-tuning Framework for Large and Small Language Models

Model ReleasesDGX agent

arXiv:2411.11707v3 Announce Type: replace-cross Abstract: By adapting Large Language Models (LLMs) to domain-specific tasks or enriching them with domain-specific knowledge, we can fully harness the c

← Previous
1…307308309310311…354
Next →