AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlog
85,115Total entries
1Added by human
85,114Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,688 results
Safety

Behavioural Analysis of Alignment Faking

DGX agent

arXiv:2605.27681v1 Announce Type: new Abstract: Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deploym

safetyarxiv-cs-ai
28 May 2026
Model Releases

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

DGX agent
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

arXiv:2605.28508v1 Announce Type: new Abstract: Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape us

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Benchmarking Fairness in Spiking Neural Networks: Data Bias, Spurious Features, and Hardware Effects

DGX agent

arXiv:2605.27407v1 Announce Type: cross Abstract: Evaluating fairness in Spiking Neural Networks (SNNs) demands rigorous benchmarks that reflect real-world complexities, yet existing assessments remai

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems

DGX agent

arXiv:2605.27492v1 Announce Type: cross Abstract: LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

DGX agent

arXiv:2605.28183v1 Announce Type: cross Abstract: We introduce the BenGER (Benchmark for German Law) dataset for evaluating LLM systems on subsumption-based legal reasoning in German law. The BenGER d

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

DGX agent

arXiv:2605.28301v1 Announce Type: new Abstract: Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond Binary Moral Judgment: Modeling Ethical Pluralism in AI

DGX agent

arXiv:2605.28707v1 Announce Type: new Abstract: Critical decision-making in socially consequential spaces is increasingly involving AI systems at varying capacities. Yet, despite the ubiquity of auton

model-releasesarxiv-cs-ai
28 May 2026
Safety

Beyond Binary: Sim-to-Real Dexterous Manipulation with Physics-Grounded Contact Representation

DGX agent

arXiv:2605.28812v1 Announce Type: cross Abstract: A primary bottleneck in contact-rich manipulation is the difficulty of collecting real-world data. Sim-to-real reinforcement learning offers a scalabl

safetyarxiv-cs-ai
28 May 2026
Model Releases

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

DGX agent

arXiv:2502.05242v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain uncl

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

DGX agent

arXiv:2509.23074v3 Announce Type: replace-cross Abstract: In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark lea

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond Motion Primitives: Behavioral Activity Recognition from Head-Mounted IMU

DGX agent

arXiv:2605.27464v1 Announce Type: cross Abstract: AR smart glasses need continuous behavioral context to offer proactive assistance, yet their most practical always-on sensor, the head-mounted Inertia

model-releasesarxiv-cs-ai
28 May 2026
Safety

BiasEdit: A Training-Free Bias-Detect-and-Edit Framework for Learning Fair Visual Classifiers

DGX agent

arXiv:2605.28450v1 Announce Type: cross Abstract: Visual data from the Web power image classifiers, which often underpin many web services, such as recommendation and content moderation. However, the

safetyarxiv-cs-ai
28 May 2026
Model Releases

BioELX: Cross-lingual Biomedical Entity Linking via Alias-based Retrieval and LLM Ranking

DGX agent

arXiv:2605.27380v1 Announce Type: cross Abstract: Cross-lingual biomedical entity linking (BEL) maps mentions in any language to unique identifiers in a biomedical knowledge base (KB), supporting clin

model-releasesarxiv-cs-ai
28 May 2026
Research

BIRDNet: Mining and Encoding Boolean Implication Knowledge Graphs as Interpretable Deep Neural Networks

DGX agent

arXiv:2605.28739v1 Announce Type: cross Abstract: Tabular data in knowledge-rich domains often carries a latent prior in the form of Boolean implication relationships (BIRs) between pairs of features.

researcharxiv-cs-ai
28 May 2026
Research

BIRDS: Characterizing and Understanding Biodiversity Impact of Large Language Model Serving

DGX agent

arXiv:2605.27480v1 Announce Type: cross Abstract: Large language model (LLM) serving creates environmental impacts beyond carbon and water, including ecosystem damage through biodiversity-related path

researcharxiv-cs-ai
28 May 2026
Model Releases

BlazeEdit: Generalist Image Editing on Mobile Devices with Image-to-Image Diffusion Models

DGX agent

arXiv:2605.28067v1 Announce Type: new Abstract: The remarkable generation quality of modern diffusion models often comes at the cost of massive parameter counts, which necessitate server-side inferenc

model-releasesarxiv-cs-ai
28 May 2026
Safety

Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking

DGX agent

arXiv:2605.28632v1 Announce Type: cross Abstract: Cryptographic watermarking is a leading defense for attributing text generated by large language models (LLMs). Existing schemes, including KGW, Unigr

safetyarxiv-cs-ai
28 May 2026
Safety

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information

DGX agent

arXiv:2605.28070v1 Announce Type: new Abstract: We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified

safetyarxiv-cs-ai
28 May 2026
Model Releases

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

DGX agent

arXiv:2605.27383v1 Announce Type: cross Abstract: Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However,

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

BuddyBench: A Privacy-Constrained Multi-Task Benchmark for Pediatric Social-Communication Personalization

DGX agent

arXiv:2605.28089v1 Announce Type: new Abstract: BuddyBench introduces a privacy-constrained multi-task benchmark for pediatric social-communication personalization. Unlike existing neurodevelopmental

model-releasesarxiv-cs-ai
28 May 2026
Tutorials

C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning

DGX agent

arXiv:2605.27860v1 Announce Type: new Abstract: Retrieval-augmented generation combined with reinforcement learning has shown promise for grounding large language models in trustworthy medical evidenc

tutorialsarxiv-cs-ai
28 May 2026
Agents

Calibrating Conservatism for Scalable Oversight

DGX agent

arXiv:2605.28807v1 Announce Type: new Abstract: Agentic AI systems capable of autonomous planning and extended environmental interaction pose a fundamental control problem: how can humans maintain mea

agentsarxiv-cs-ai
28 May 2026
Research

CaMBRAIN: Real-time, Continuous EEG Inference with Causal State Space Models

DGX agent

arXiv:2605.28792v1 Announce Type: new Abstract: Electroencephalography (EEG) is a critical, non-invasive method to monitor electrical brain activity. EEGs can span anywhere from a couple seconds to mu

researcharxiv-cs-ai
28 May 2026
Research

Can I Have Your Order? Monte-Carlo Tree Search for Slot Filling Ordering in Diffusion Language Models

DGX agent

arXiv:2602.12586v2 Announce Type: replace Abstract: While plan-and-infill decoding in Masked Diffusion Models (MDMs) shows promise for mathematical and code reasoning, performance remains highly sensi

researcharxiv-cs-ai
28 May 2026
Research

Can Quantum Federated Learning Withstand Circuit-Level Backdoors?

DGX agent

arXiv:2605.27416v1 Announce Type: cross Abstract: Quantum Federated Learning (QFL) inherits the core vulnerability of federated optimization to malicious clients, while also introducing an attack surf

researcharxiv-cs-ai
28 May 2026
Model Releases

Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought

DGX agent

arXiv:2605.27764v1 Announce Type: cross Abstract: Recent segmentation models couple large language models (LLMs) with mask decoders to ground complex language expressions into masks, yet their instruc

model-releasesarxiv-cs-ai
28 May 2026
Safety

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

DGX agent

arXiv:2603.22335v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior dis

safetyarxiv-cs-ai
28 May 2026
Tutorials

Checking Fact with Better Retrieval: Dynamic Contrastive Learning for Evidence Retrieval

DGX agent

arXiv:2605.27449v1 Announce Type: cross Abstract: In the field of multimodal fact checking, the accuracy of retrieving evidence from different modalities has a significant impact on the downstream cla

tutorialsarxiv-cs-ai
28 May 2026
Model Releases

ChildEval: When large language models meet children's personalities

DGX agent

arXiv:2605.27805v1 Announce Type: cross Abstract: While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-spec

model-releasesarxiv-cs-ai
28 May 2026
Agents

CircuitLM: A Multi-Agent LLM-Aided Design Framework for Generating Circuit Schematics from Natural Language Prompts

DGX agent

arXiv:2601.04505v3 Announce Type: replace Abstract: Generating accurate circuit schematics from high-level natural language descriptions remains a persistent challenge in electronic design automation

agentsarxiv-cs-ai
28 May 2026
Model Releases

CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text

DGX agent

arXiv:2605.27700v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while contain

model-releasesarxiv-cs-ai
28 May 2026
Local Ai

CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models

DGX agent

arXiv:2605.28115v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face severe memory and latency bottlenecks due to high-resolution visual tokens. While current token reduction methods the

local-aiarxiv-cs-ai
28 May 2026
Local Ai

CLANE: Continual Learning of Actions on Neuromorphic Hardware from Event Cameras

DGX agent

arXiv:2605.28387v1 Announce Type: cross Abstract: Recognizing and continuously learning novel human actions without forgetting prior classes is a requirement for emerging AR/VR and robotics applicatio

local-aiarxiv-cs-ai
28 May 2026
Research

Clark Hash: Stateless Sparse Johnson-Lindenstrauss Quantization for Neural Embeddings

DGX agent

arXiv:2605.28034v1 Announce Type: new Abstract: Clark Hash is a small method for storing neural embeddings in less space. It normalizes each database vector, applies a deterministic sparse signed John

researcharxiv-cs-ai
28 May 2026
Research

Clinical Validation of the Melanoscope AI Mobile Dermoscopy Clinical Decision Support System

DGX agent

arXiv:2605.27561v1 Announce Type: cross Abstract: Introduction. Early detection of malignant skin lesions is critical for prognosis, yet dermatologist shortages in Russian regions limit screening cove

researcharxiv-cs-ai
28 May 2026
Model Releases

Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code

DGX agent

arXiv:2603.24631v2 Announce Type: replace-cross Abstract: Code agents resolve 65-70% of SWE-bench Verified issues, but Pass@1 cannot tell us why the rest fail, and, as we show, capable-model failures

model-releasesarxiv-cs-ai
28 May 2026
Safety

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

DGX agent

arXiv:2602.15198v2 Announce Type: replace-cross Abstract: Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperativ

safetyarxiv-cs-ai
28 May 2026
Model Releases

Comparative Analysis of Liquid Neural Networks and LSTM for Sequential Pattern Recognition: Robustness, Efficiency, and Clinical Utility

DGX agent

arXiv:2605.27467v1 Announce Type: cross Abstract: Traditional Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) units operate on discrete time steps, often failing to capture the flui

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback

DGX agent

arXiv:2605.28010v1 Announce Type: new Abstract: Self-evolving large language models (LLMs) learn by generating their own training tasks and solutions, reducing reliance on human-curated supervision. H

model-releasesarxiv-cs-ai
28 May 2026
Research

Constrained Auto-Bidding via Generative Response Modeling

DGX agent

arXiv:2605.27811v1 Announce Type: new Abstract: Auto-bidding systems aim to maximize advertiser value over long horizons under budget constraints and ratio targets such as cost-per-acquisition, yet fu

researcharxiv-cs-ai
28 May 2026
Model Releases

Continual Model Routing in Evolving Model Hubs

DGX agent

arXiv:2605.28577v1 Announce Type: new Abstract: AI model hubs provide access to a rapidly growing collection of powerful pre-trained models, enabling off-the-shelf mixture-of-experts systems with diff

model-releasesarxiv-cs-ai
28 May 2026
Agents

COOP^2: Defining, Observing, and Repairing Cooperation in LLM Multi-Agent Systems

DGX agent

arXiv:2603.00349v2 Announce Type: replace Abstract: Many complex tasks require extended effort, diverse capabilities, or coordinated actions beyond what a single agent can provide. However, simply add

agentsarxiv-cs-ai
28 May 2026
Research

CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning

DGX agent

arXiv:2605.28742v1 Announce Type: new Abstract: Language models can use verifiable rewards to improve at a wide variety of reasoning tasks. However, both parametric (e.g. RLVR) and non-parametric (e.g

researcharxiv-cs-ai
28 May 2026
Safety

COTTA: Context-Aware Transfer Adaptation for Trajectory Prediction in Autonomous Driving

DGX agent

arXiv:2604.00402v2 Announce Type: replace-cross Abstract: Developing robust models to accurately predict the trajectories of surrounding agents is fundamental to autonomous driving safety. However, mo

safetyarxiv-cs-ai
28 May 2026
Safety

CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders

DGX agent

arXiv:2604.01604v2 Announce Type: replace Abstract: While modern LLMs are aligned to refuse harmful requests, it is essential to understand the underlying mechanistic basis of this refusal behavior fo

safetyarxiv-cs-ai
28 May 2026
Safety

Cross-Entropy Games and Frost Training

DGX agent

arXiv:2605.27701v1 Announce Type: new Abstract: We present Frost Training, a method for improving Monte Carlo-based policy optimization for a large family of LLM-as-a-judge tasks called Cross-Entropy

safetyarxiv-cs-ai
28 May 2026
Research

CubePart: An Open-Vocabulary Part-Controllable 3D Generator

DGX agent

arXiv:2605.28763v1 Announce Type: new Abstract: Interactive 3D assets used in games and simulation are typically decomposed into specific semantic parts to support animation, physics, and scripted beh

researcharxiv-cs-ai
28 May 2026
Model Releases

Cultural Binding Heads in Language Models

DGX agent

arXiv:2605.28543v1 Announce Type: new Abstract: LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness. Usin

model-releasesarxiv-cs-ai
28 May 2026
← Previous
1…250251252253254…452
Next →