AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,271
  • Agents7,543
  • Applications5,408
  • Concepts5
  • Hardware1,828
  • Industry6,162
  • Local Ai4,927
  • Model Releases23,818
  • Research20,122
  • Safety13,367
  • Syntheses17
  • Tools1,674
  • Tutorials3,400

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,271
  • Agents7,543
  • Applications5,408
  • Concepts5
  • Hardware1,828
  • Industry6,162
  • Local Ai4,927
  • Model Releases23,818
  • Research20,122
  • Safety13,367
  • Syntheses17
  • Tools1,674
  • Tutorials3,400

Source
Human
88,271Total entries
1Added by human
88,270Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
62,897 results
29 May 2026

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures

SafetyDGX agent

arXiv:2605.29629v1 Announce Type: new Abstract: Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not ho

Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning

SafetyDGX agent

arXiv:2605.29414v1 Announce Type: cross Abstract: Recent studies have shown that code-switching data (CSD), in which multiple languages are mixed within the same context, can improve cross-lingual tra

Beyond Consensus: Trace-Level Synthesis in Mixture of Agents

AgentsDGX agent

arXiv:2605.29116v1 Announce Type: new Abstract: When multiple LLM agents solve the same problem, standard practice compresses each agent's reasoning into a majority vote or layered synthesis, treating

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

Model ReleasesDGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

Beyond MSE: Improving Precipitation Nowcasting with Multi-Quantile Regression

ResearchDGX agent

arXiv:2605.30122v1 Announce Type: cross Abstract: Deep-learning precipitation nowcasting models are often optimized using pointwise losses such as mean squared error or mean absolute error, which can

Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR

ResearchDGX agent

arXiv:2602.12642v2 Announce Type: replace-cross Abstract: Reward-maximizing RL methods have shown to be capable of enhancing the reasoning performance of LLMs, but often lead to reduced generation div

Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization

Model ReleasesDGX agent

arXiv:2605.28969v1 Announce Type: cross Abstract: If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how f

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

SafetyDGX agent

arXiv:2605.29697v1 Announce Type: new Abstract: In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward

Beyond Transcripts: A Renewed Perspective on Audio Chaptering

ResearchDGX agent

arXiv:2602.08979v2 Announce Type: replace-cross Abstract: Audio chaptering, the task of segmenting long-form audio into coherent sections, is increasingly important for navigating podcasts, lectures,

BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models

TutorialsDGX agent

arXiv:2512.00283v3 Announce Type: replace-cross Abstract: Foundation models have revolutionized various fields such as natural language processing (NLP) and computer vision (CV). While efforts have be

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2605.30162v1 Announce Type: new Abstract: Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model

BitC-3DGS: High-Capacity 3D Gaussian Splatting Watermarking via Bit Compression

ResearchDGX agent

arXiv:2605.29583v1 Announce Type: new Abstract: High-capacity watermarking is necessary for 3D Gaussian Splatting (3DGS) assets to embed rich information (e.g., ownership, provenance, and authenticati

BitTP: The Lightweight Trajectory Prediction Model with BitLLM for Edge-Devices

AgentsDGX agent

arXiv:2605.29705v1 Announce Type: new Abstract: Trajectory prediction is a fundamental task for autonomous systems, requiring complex reasoning about multi-agent interactions and intents. Large langua

BlockBatch: Multi-Scale Consensus Decoding for Efficient Diffusion Language Model Inference

Local AiDGX agent

arXiv:2605.29233v1 Announce Type: cross Abstract: Diffusion language models (dLLMs) generate text by iteratively denoising multiple token positions in parallel, offering an attractive alternative to s

Boosting Image Quality Assessment Performance: Unsupervised Score Fusion by Deep Maximum a Posteriori Estimation

ResearchDGX agent

arXiv:2605.30269v1 Announce Type: new Abstract: Over the past decades, numerous Image Quality Assessment (IQA) models have emerged, aiming to predict the perceptual quality of images. However, individ

Boosting Zero-Shot 3D Style Transfer with 2D Pre-trained Priors

ResearchDGX agent

arXiv:2605.30065v1 Announce Type: new Abstract: In this work, we focus on zero-shot 3D style transfer that can generate multi-view consistent stylized views of the 3D scene given an arbitrary style im

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models

SafetyDGX agent

arXiv:2605.30226v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulat

Bosses, Kings, and the Commons: Cooperation Under Power Asymmetry in LLM Societies

AgentsDGX agent

arXiv:2605.29062v1 Announce Type: new Abstract: Communities can sustainably manage shared resources (commons) through self-governance and cooperative norms, a central finding of Ostrom's theory of sel

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base

Model ReleasesDGX agent

arXiv:2605.29379v1 Announce Type: new Abstract: We present BrahmicTokenizer-131K, a 131,072-vocabulary byte-level BPE tokenizer that closes the Brahmic compression gap at the 131K-vocabulary class whi

Brain-IT-VQA: From Brain Signals to Answers

Model ReleasesDGX agent

arXiv:2605.29588v1 Announce Type: cross Abstract: Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-

Bridge-RAG: An Abstract Bridge Tree Based Retrieval Augmented Generation Algorithm

ResearchDGX agent

arXiv:2603.26668v2 Announce Type: replace-cross Abstract: As an important paradigm for enhancing the generation quality of Large Language Models (LLMs), retrieval-augmented generation (RAG) faces the

Bridging Chemists and AI: An Expert-Augmented Framework for Interpretable Route Evaluation

ResearchDGX agent

arXiv:2605.29108v1 Announce Type: new Abstract: Selecting efficient multi-step synthetic routes is a central challenge in organic synthesis, particularly in medicinal and process chemistry, where rout

Bridging Functional and Representational Similarity via Usable Information

ResearchDGX agent

arXiv:2601.21568v2 Announce Type: replace Abstract: We present a unified framework for quantifying the similarity between representations through the lens of extit{usable} information, offering a rigo

Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models

Model ReleasesDGX agent

arXiv:2601.01162v3 Announce Type: replace-cross Abstract: Qualitative data are widespread in domains such as healthcare, marketing, and bioinformatics, where clustering offers a fundamental tool for p

Bridging the Sim-to-Real Gap in Reinforcement Learning-Based Industrial Dispatching through Execution Semantics

SafetyDGX agent

arXiv:2605.29078v1 Announce Type: new Abstract: Event-driven scheduling policies are increasingly deployed in industrial environments, where decisions are made under asynchronous and partially observe

Building and Road Recognition in Dense Urban Informal Settlements: A Dataset and Benchmark

Model ReleasesDGX agent

arXiv:2605.29856v1 Announce Type: new Abstract: As a widespread form of informal settlements, urban villages present significant challenges for sustainable urban development and governance. Precise ma

BuilDyn: Excitation-Driven Data Generation for Building Thermal Dynamics Modeling and Control

ApplicationsDGX agent

arXiv:2605.29849v1 Announce Type: cross Abstract: Machine learning (ML) is increasingly used for data-driven modeling of buildings to enable downstream tasks such as fault detection and diagnosis, and

BullingerDB: A Dataset for Handwritten Text Recognition and Writer Retrieval

Model ReleasesDGX agent

arXiv:2605.30235v1 Announce Type: new Abstract: We present BullingerDB, a large-scale benchmark dataset for historical document analysis based on the correspondence of Heinrich Bullinger (1504-1575).

CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive Control

HardwareDGX agent

arXiv:2605.29155v1 Announce Type: cross Abstract: In the literature, actor-critic model predictive control (AC-MPC) integrates MPC with reinforcement learning to enable high-performance control of com

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

Model ReleasesDGX agent

arXiv:2605.30188v1 Announce Type: cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibr

Calibrating Generative Models to Distributional Constraints

ResearchDGX agent

arXiv:2510.10020v4 Announce Type: replace-cross Abstract: Generative models frequently suffer miscalibration, wherein statistics of the sampling distribution, such as the fraction of generations in a

Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

SafetyDGX agent

arXiv:2601.08064v2 Announce Type: replace Abstract: Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing eval

CamC2V: Context-aware Controllable Video Generation

TutorialsDGX agent

arXiv:2504.06022v3 Announce Type: replace Abstract: Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditi

Can AI Weather Models Predict Beyond Two Weeks? A Quantitative Benchmark and Analysis of Long Rollouts

Model ReleasesDGX agent

arXiv:2605.30184v1 Announce Type: new Abstract: While AI weather models excel at short-to-medium range forecasts (up to 15 days), they frequently suffer from ill-defined 'instabilities' when rolled ou

CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation

ResearchDGX agent

arXiv:2605.29316v1 Announce Type: new Abstract: Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing met

Casual as an Anchor: Resolving Supervision Misalignment in Formality Transfer Dataset

Model ReleasesDGX agent

arXiv:2605.29365v1 Announce Type: new Abstract: Formality transfer is commonly framed as a symmetric bidirectional task between informal and formal registers. We argue that this framing conceals a sup

Catalyst-Agent: Autonomous heterogeneous catalyst screening with an LLM Agent

AgentsDGX agent

arXiv:2603.01311v2 Announce Type: replace Abstract: The discovery of novel catalysts tailored for particular applications is a major challenge for the twenty-first century. Traditional methods for thi

Causal Intelligence for Constraint-Aware Intervention Design to Induce State Transitions

TutorialsDGX agent

arXiv:2605.29008v1 Announce Type: new Abstract: Driving a system from one state to another through targeted interventions is a fundamental challenge in science, yet most predictive models offer limite

Causal Interventions on Continuous Variables: A Case Study on Verb Bias in Steering Vectors for In-Context Learning

SafetyDGX agent

arXiv:2605.29971v1 Announce Type: new Abstract: Causal interventions in language model representations have largely targeted discrete features, like grammatical number. However, language models must a

Causal-JEPA: Learning World Models through Object-Level Latent Masking

SafetyDGX agent

arXiv:2602.11389v2 Announce Type: replace Abstract: World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a u

Causal Label Recovery in Payment Networks

ResearchDGX agent

arXiv:2605.29272v1 Announce Type: cross Abstract: Fraud detection models in payment networks train on chargeback labels that are systematically biased. Every label must survive three sequential gates:

CB-SLICE: Concept-Based Interpretable Error Slice Discovery

SafetyDGX agent

arXiv:2605.29836v1 Announce Type: cross Abstract: Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Id

CCS: Clinical Consensus Selection for Radiology Report Generation

ResearchDGX agent

arXiv:2605.30131v1 Announce Type: new Abstract: Radiology report generation (RRG) is commonly formulated as a single-path generation task, where a multimodal large language model (MLLM) produces one d

Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing

ResearchDGX agent

arXiv:2605.29809v1 Announce Type: cross Abstract: Large-scale text-to-image (T2I) diffusion models have enabled unprecedented creative applications, but their unauthorized use has raised serious intel

Certified Causal Defense with Generalizable Robustness

Model ReleasesDGX agent

arXiv:2408.15451v3 Announce Type: replace Abstract: While machine learning models have proven effective across various scenarios, it is widely acknowledged that many models are vulnerable to adversari

Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk

SafetyDGX agent

arXiv:2605.29788v1 Announce Type: new Abstract: Critical sequential decisions are rarely single-timescale: a strategic decision causally shapes the context in which every subsequent tactical choice is

Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences

Model ReleasesDGX agent

arXiv:2605.30100v1 Announce Type: new Abstract: World models require state tracking, which is the ability to maintain a correct latent state across action sequences. Existing benchmarks are often synt

Ciphera: A Decentralised Biometric Identity Framework

ApplicationsDGX agent

arXiv:2605.29868v1 Announce Type: cross Abstract: Centralised biometric identity systems expose users to single points of failure, opaque verification processes, and irreversible biometric compromise.

Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering

Model ReleasesDGX agent

arXiv:2605.29742v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority

City-Mesh3R: Simulation-Ready City-Scale 3D Mesh Reconstruction from Multi-View Images

ResearchDGX agent

arXiv:2605.30310v1 Announce Type: cross Abstract: City-scale 3D surface reconstruction from multiview images for downstream 3D simulation, poses highly challenging problems due to the scale and comple

CityGen: Structure-Guided City-Style Synthesis for Cross-City Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.29935v1 Announce Type: cross Abstract: Autonomous driving systems are commonly trained and evaluated within limited geographic regions, which hinders their scalability when deployed in new

Classification of non-analyzable word types in web documents to implement an effective Korean e-learning system

Local AiDGX agent

arXiv:2605.29638v1 Announce Type: new Abstract: E-learning systems should deliver contents that reflect various phenomena of the language as it is used. In addition to formal Korean, e-learning system

CLUBench: A Clustering Benchmark

Model ReleasesDGX agent

arXiv:2605.29933v1 Announce Type: new Abstract: Clustering is a fundamental problem in data science with a long-standing research history, yielding numerous insightful algorithms. Despite this progres

Cluster-Level Attention-Guided Parallel Decoding for Masked Diffusion Language Models

ResearchDGX agent

arXiv:2605.29607v1 Announce Type: new Abstract: Masked diffusion language models (MDLMs) enable parallel decoding by predicting all masked positions at each denoising step, yet existing training-free

Coarse-Grained Boltzmann Generators

ResearchDGX agent

arXiv:2602.10637v2 Announce Type: replace Abstract: Sampling equilibrium molecular configurations from the Boltzmann distribution is a longstanding challenge. Boltzmann Generators (BGs) address this b

Code-QA-Bench: Separating Code Reasoning from Documentation Memorization in Repository-Level QA

AgentsDGX agent

arXiv:2605.29277v1 Announce Type: cross Abstract: We present Code-QA-Bench, a fully automated framework for synthesizing repository-level code understanding benchmarks that separates genuine code comp

CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization

Model ReleasesDGX agent

arXiv:2510.14150v5 Announce Type: replace Abstract: We introduce CodeEvolve, an open-source framework that couples large language models with island-based evolutionary search for end-to-end algorithmi

Cognitive Loop of Thought: Reversible Hierarchical Markov Chain for Efficient Mathematical Reasoning

ResearchDGX agent

arXiv:2604.06805v2 Announce Type: replace Abstract: Multi-step Chain-of-Thought (CoT) has significantly advanced the mathematical reasoning capabilities of LLMs by leveraging explicit reasoning steps.

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning

ResearchDGX agent

arXiv:2605.29602v1 Announce Type: new Abstract: Multi-modal Retrieval-Augmented Generation (MMRAG) has emerged as a powerful paradigm for enhancing Multimodal Large Language Models in knowledge-intens

CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval

ResearchDGX agent

arXiv:2605.29271v1 Announce Type: new Abstract: Tool retrieval over large API catalogs is a core bottleneck for LLM agents: user queries arrive in colloquial, often underspecified language, while the

← Previous
1…551552553554555…1049
Next →