AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
29 May 2026

AgentSchool: An LLM-Powered Multi-Agent Simulation for Education

AgentsDGX agent

arXiv:2605.30144v1 Announce Type: new Abstract: Despite the rapid deployment of LLMs into classrooms, validating educational AI remains uniquely intractable: interventions act on developing learners w

Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents

SafetyDGX agent

arXiv:2605.29910v1 Announce Type: cross Abstract: Consensus protocols form the backbone of distributed systems and blockchains, where implementation bugs can cause data corruption and financial losses

AIRGuard: Guarding Agent Actions with Runtime Authority Control

SafetyDGX agent

arXiv:2605.28914v1 Announce Type: cross Abstract: Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model C


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization

Model ReleasesDGX agent

arXiv:2605.29396v1 Announce Type: new Abstract: Safety alignment for large language models (LLMs) aims to reduce harmful or unsafe behavior while preserving general utility. However, recent findings r

Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models

Model ReleasesDGX agent

arXiv:2605.30038v1 Announce Type: cross Abstract: Diffusion models generate highly realistic images but often struggle with precise text-image alignment. While recent post-training methods improve ali

AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing

SafetyDGX agent

arXiv:2605.29434v1 Announce Type: cross Abstract: Existing sentence-level watermarking methods enhance robustness to paraphrasing by anchoring watermarks in sentence semantics. However, their prefix-b

An accuracy-aware extension to LRP-based pruning for CNNs to prevent cascading accuracy degradation in data-scarce transfer learning

ResearchDGX agent

arXiv:2511.10861v3 Announce Type: replace-cross Abstract: Convolutional Neural Networks (CNNs) pre-trained on large-scale datasets such as ImageNet are widely used as feature extractors to construct h

Anchorless Diversification for Parallel LLM Ideation

ResearchDGX agent

arXiv:2605.30150v1 Announce Type: new Abstract: LLMs are increasingly used to generate candidate-idea pools for creative tasks where broad exploration is valuable. Parallel inference can be attractive

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

ResearchDGX agent

arXiv:2605.29488v1 Announce Type: cross Abstract: Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are

Approximate Proportionality in Online Fair Division

ResearchDGX agent

arXiv:2508.03253v2 Announce Type: replace-cross Abstract: We study the online fair division problem, where indivisible goods arrive sequentially and must be allocated immediately and irrevocably. Prio

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark

Model ReleasesDGX agent

arXiv:2605.29400v1 Announce Type: new Abstract: We benchmark three supervised fine-tuned models against frontier zero-shot baselines on a 661-row held-out slice of PiSAR (Persona, intent, Screen, Acti

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

ResearchDGX agent

arXiv:2605.30311v1 Announce Type: cross Abstract: Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visu

Are LLMs Socially Adaptive? Contrasting Belief Evolution in Large Language Models and Humans

Model ReleasesDGX agent

arXiv:2410.10398v3 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly engage in complex social interactions, ensuring that their behaviors align with human ethical pri

Aryabhata 2: Scaling Reinforcement Learning for Advanced STEM Reasoning

ApplicationsDGX agent

arXiv:2605.28829v1 Announce Type: cross Abstract: Competitive STEM examinations such as JEE and NEET require multi-step symbolic reasoning, precise numerical computation, and deep conceptual understan

Assessing Dutch Syllabification Algorithms and Improving Accuracy by Combining Phonetic and Orthographic Information through Deep Learning

ResearchDGX agent

arXiv:2605.28834v1 Announce Type: cross Abstract: Syllabification describes the task of dividing words into syllables. Due to many rules and exceptions, training an algorithm to perform syllabificatio

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials

Model ReleasesDGX agent

arXiv:2510.04704v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promising potential in scientific research, enabling tasks ranging from knowledge retrieval to propert

AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence

Model ReleasesDGX agent

arXiv:2605.21739v2 Announce Type: replace Abstract: Emotional intelligence (EI), the ability to perceive, understand, and respond appropriately to others' emotional states, is central to human communi

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

SafetyDGX agent

arXiv:2605.30031v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) expand jailbreak risks from token-level prompting to the full speech perception-to-reasoning pipeline, where unsaf

Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency

SafetyDGX agent

arXiv:2605.30208v1 Announce Type: cross Abstract: AI-assisted coding tools have altered software production. At Meta, significant lines of code per human-landed diff grew by 105.9% year over year and

AutoSizer: Automatic Sizing of Analog and Mixed-Signal Circuits via Large Language Model (LLM) Agents

Model ReleasesDGX agent

arXiv:2602.02849v2 Announce Type: replace Abstract: The design of Analog and Mixed-Signal (AMS) integrated circuits remains heavily reliant on expert knowledge, with transistor sizing a major bottlene

Balancing Multimodal Learning through Label Space Reshaping

Model ReleasesDGX agent

arXiv:2605.28869v1 Announce Type: cross Abstract: Multimodal learning often suffers from modality imbalance, where modalities that converge faster dominate optimization while others remain undertraine

Battery-Sim-Agent: Leveraging LLM-Agent for Inverse Battery Parameter Estimation

Model ReleasesDGX agent

arXiv:2605.29560v1 Announce Type: new Abstract: Parameterizing high-fidelity 'digital twins' of batteries is a critical yet challenging inverse problem that hinders the pace of battery innovation. Pre

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

SafetyDGX agent

arXiv:2605.28994v1 Announce Type: new Abstract: AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable.

Before the Shutter: Aesthetic and Actionable Portrait Photography Planning in 3D Scenes

ApplicationsDGX agent

arXiv:2605.30318v1 Announce Type: cross Abstract: Portrait photography is largely decided before the shutter opens: the subject's pose, the camera configuration, and the lighting devices must be coord

Behavior-Aware Auxiliary Corrections for Off-Policy Temporal-Difference Prediction

Local AiDGX agent

arXiv:2605.28855v1 Announce Type: new Abstract: Temporal-difference learning with function approximation can be unstable under off-policy sampling. TDC stabilizes off-policy TD through an auxiliary co

Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction

SafetyDGX agent

arXiv:2605.28849v1 Announce Type: new Abstract: Gradient temporal-difference methods provide stable off-policy prediction with linear function approximation, but their practical performance is strongl

Benchmarking at the Edge of Comprehension

Local AiDGX agent

arXiv:2602.14307v3 Announce Type: replace Abstract: As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

Model ReleasesDGX agent

arXiv:2605.29462v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified i

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting

Model ReleasesDGX agent

arXiv:2509.23571v3 Announce Type: replace-cross Abstract: As cyber threats continue to grow in scale and sophistication, blue team defenders increasingly require advanced tools to proactively detect a

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

Model ReleasesDGX agent

arXiv:2605.28830v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a c

Benchmarking Positional Encoding Strategies for Transformer-Based EEG Foundation Models

Model ReleasesDGX agent

arXiv:2605.29754v1 Announce Type: new Abstract: Electroencephalography (EEG) is a widely used non-invasive technique for measuring brain activity in brain-computer interface (BCI) applications. Superv

BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

Model ReleasesDGX agent

arXiv:2605.29225v1 Announce Type: new Abstract: Self-evolving agents improve over time by reflecting on past failures, but existing evaluation is limited in two ways: it measures only task scores, lea

Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-grounded Post-extraction Correction

ResearchDGX agent

arXiv:2605.29168v1 Announce Type: new Abstract: Question answering (QA) is a core challenge in AI, particularly for complex queries requiring multi-hop reasoning across documents, or symbolic operatio

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning

ResearchDGX agent

arXiv:2605.30231v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-ans

Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?

ResearchDGX agent

arXiv:2510.16060v2 Announce Type: replace-cross Abstract: The recent development of foundation models for time series data has generated considerable interest in using such models across a variety of

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures

SafetyDGX agent

arXiv:2605.29629v1 Announce Type: new Abstract: Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not ho

Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning

SafetyDGX agent

arXiv:2605.29414v1 Announce Type: cross Abstract: Recent studies have shown that code-switching data (CSD), in which multiple languages are mixed within the same context, can improve cross-lingual tra

Beyond Consensus: Trace-Level Synthesis in Mixture of Agents

AgentsDGX agent

arXiv:2605.29116v1 Announce Type: new Abstract: When multiple LLM agents solve the same problem, standard practice compresses each agent's reasoning into a majority vote or layered synthesis, treating

Beyond MSE: Improving Precipitation Nowcasting with Multi-Quantile Regression

ResearchDGX agent

arXiv:2605.30122v1 Announce Type: cross Abstract: Deep-learning precipitation nowcasting models are often optimized using pointwise losses such as mean squared error or mean absolute error, which can

Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR

ResearchDGX agent

arXiv:2602.12642v2 Announce Type: replace-cross Abstract: Reward-maximizing RL methods have shown to be capable of enhancing the reasoning performance of LLMs, but often lead to reduced generation div

Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization

Model ReleasesDGX agent

arXiv:2605.28969v1 Announce Type: cross Abstract: If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how f

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

SafetyDGX agent

arXiv:2605.29697v1 Announce Type: new Abstract: In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward

BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models

TutorialsDGX agent

arXiv:2512.00283v3 Announce Type: replace-cross Abstract: Foundation models have revolutionized various fields such as natural language processing (NLP) and computer vision (CV). While efforts have be

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2605.30162v1 Announce Type: new Abstract: Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model

BitTP: The Lightweight Trajectory Prediction Model with BitLLM for Edge-Devices

AgentsDGX agent

arXiv:2605.29705v1 Announce Type: new Abstract: Trajectory prediction is a fundamental task for autonomous systems, requiring complex reasoning about multi-agent interactions and intents. Large langua

BlockBatch: Multi-Scale Consensus Decoding for Efficient Diffusion Language Model Inference

Local AiDGX agent

arXiv:2605.29233v1 Announce Type: cross Abstract: Diffusion language models (dLLMs) generate text by iteratively denoising multiple token positions in parallel, offering an attractive alternative to s

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models

SafetyDGX agent

arXiv:2605.30226v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulat

Brain-IT-VQA: From Brain Signals to Answers

Model ReleasesDGX agent

arXiv:2605.29588v1 Announce Type: cross Abstract: Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-

Bridge-RAG: An Abstract Bridge Tree Based Retrieval Augmented Generation Algorithm

ResearchDGX agent

arXiv:2603.26668v2 Announce Type: replace-cross Abstract: As an important paradigm for enhancing the generation quality of Large Language Models (LLMs), retrieval-augmented generation (RAG) faces the

Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models

Model ReleasesDGX agent

arXiv:2601.01162v3 Announce Type: replace-cross Abstract: Qualitative data are widespread in domains such as healthcare, marketing, and bioinformatics, where clustering offers a fundamental tool for p

Bridging the Sim-to-Real Gap in Reinforcement Learning-Based Industrial Dispatching through Execution Semantics

SafetyDGX agent

arXiv:2605.29078v1 Announce Type: new Abstract: Event-driven scheduling policies are increasingly deployed in industrial environments, where decisions are made under asynchronous and partially observe

CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive Control

HardwareDGX agent

arXiv:2605.29155v1 Announce Type: cross Abstract: In the literature, actor-critic model predictive control (AC-MPC) integrates MPC with reinforcement learning to enable high-performance control of com

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

Model ReleasesDGX agent

arXiv:2605.30188v1 Announce Type: cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibr

Causal-JEPA: Learning World Models through Object-Level Latent Masking

SafetyDGX agent

arXiv:2602.11389v2 Announce Type: replace Abstract: World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a u

Causal Label Recovery in Payment Networks

ResearchDGX agent

arXiv:2605.29272v1 Announce Type: cross Abstract: Fraud detection models in payment networks train on chargeback labels that are systematically biased. Every label must survive three sequential gates:

CB-SLICE: Concept-Based Interpretable Error Slice Discovery

SafetyDGX agent

arXiv:2605.29836v1 Announce Type: cross Abstract: Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Id

Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk

SafetyDGX agent

arXiv:2605.29788v1 Announce Type: new Abstract: Critical sequential decisions are rarely single-timescale: a strategic decision causally shapes the context in which every subsequent tactical choice is

Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering

Model ReleasesDGX agent

arXiv:2605.29742v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority

City-Mesh3R: Simulation-Ready City-Scale 3D Mesh Reconstruction from Multi-View Images

ResearchDGX agent

arXiv:2605.30310v1 Announce Type: cross Abstract: City-scale 3D surface reconstruction from multiview images for downstream 3D simulation, poses highly challenging problems due to the scale and comple

CityGen: Structure-Guided City-Style Synthesis for Cross-City Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.29935v1 Announce Type: cross Abstract: Autonomous driving systems are commonly trained and evaluated within limited geographic regions, which hinders their scalability when deployed in new

← Previous
1…188189190191192…358
Next →