AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
Model Releases

Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models

DGX agent

arXiv:2503.01829v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) demonstrate persuasive capabilities that rival human-level persuasion. While these capabilities can be used for s

model-releasesarxiv-cs-ai
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

PetroBench: A Benchmark for Large Language Models in Petroleum Engineering

DGX agent

arXiv:2605.28032v1 Announce Type: new Abstract: Large Language Models are increasingly applied in the petroleum industry, highlighting the need for a domain-specific evaluation framework. This study d

model-releasesarxiv-cs-ai
28 May 2026
Safety

Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains

DGX agent

arXiv:2605.28345v1 Announce Type: new Abstract: Progress in Prognostics and Health Management (PHM) is hindered by the lack of standardized and reusable evaluation practices across tasks, datasets, an

safetyarxiv-cs-ai
28 May 2026
Agents

PIRS: Physics-Informed Reward Shaping for SAC-Based Building Energy Management

DGX agent

arXiv:2605.28232v1 Announce Type: new Abstract: Occupant comfort and grid-aware energy efficiency are competing objectives whose joint optimization depends critically on how reward functions are speci

agentsarxiv-cs-ai
28 May 2026
Agents

Plan Before Search: Search Agents Need Plan

DGX agent

arXiv:2605.28354v1 Announce Type: new Abstract: Training large language models as retrieval-augmented reasoning agents typically combines reinforcement learning with an SFT cold start distilled from a

agentsarxiv-cs-ai
28 May 2026
Research

Planning a Community Approach to Diabetes Care in Low- and Middle-Income Countries Using Optimization

DGX agent

arXiv:2305.06426v2 Announce Type: replace Abstract: Diabetes is a global health priority, especially in low- and-middle-income countries, where over 50% of premature deaths are attributed to high bloo

researcharxiv-cs-ai
28 May 2026
Model Releases

Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

DGX agent

arXiv:2605.28201v1 Announce Type: new Abstract: Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into ext

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

DGX agent

arXiv:2605.27887v1 Announce Type: new Abstract: LLMs have shown strong performance across diverse financial tasks, yet portfolio management (PM), a critical financial decision-making task, remains poo

model-releasesarxiv-cs-ai
28 May 2026
Safety

Position: Retire the 'Positive Backdoor' Label -- Secret Alignment Requires Strict and Systematic Evaluation

DGX agent

arXiv:2605.28597v1 Announce Type: cross Abstract: This position paper argues that the AI/ML community should stop overclaiming and retire the label 'positive backdoor,' and instead treat trigger-activ

safetyarxiv-cs-ai
28 May 2026
Research

Preference-Shaped Expected Hypervolume and R2 Improvement: Exact Computation and Monotonicity

DGX agent

arXiv:2605.28746v1 Announce Type: cross Abstract: This paper studies preference-shaped expected improvement criteria for Bayesian multiobjective optimization. We consider two indicator families which

researcharxiv-cs-ai
28 May 2026
Research

Prefix-Safe Bayesian Belief Tracking for LLM Reasoning Reliability:Separating Calibration from Ranking

DGX agent

arXiv:2605.27712v1 Announce Type: new Abstract: Long reasoning traces need reliability estimates before final answers are known. We study prefix-conditioned eventual-success estimation, P(y=1 mid o_{1

researcharxiv-cs-ai
28 May 2026
Model Releases

Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

DGX agent

arXiv:2605.27958v1 Announce Type: cross Abstract: Linear probes trained on LLM activations are increasingly proposed as deception-detection metrics, yet report AUROC exceeding 0.96 on clean benchmarks

model-releasesarxiv-cs-ai
28 May 2026
Safety

Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

DGX agent

arXiv:2602.01745v2 Announce Type: replace-cross Abstract: Token-level reweighting is a simple yet effective mechanism for controlling supervised fine-tuning, but common indicators are largely one-dime

safetyarxiv-cs-ai
28 May 2026
Model Releases

Probing for Knowledge Attribution in Large Language Models

DGX agent

arXiv:2602.22787v2 Announce Type: replace-cross Abstract: Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, w

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit

DGX agent

arXiv:2605.27439v1 Announce Type: cross Abstract: AI assistants like ChatGPT and Claude are recommendation engines, not search engines: they answer commercial queries by directly nominating brands rat

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement

DGX agent

arXiv:2605.28360v1 Announce Type: new Abstract: Automatic prompt optimization (APO) has driven significant gains in LLM-based agentic workflows. However, existing methods treat each task's prompt as a

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting

DGX agent

arXiv:2605.28066v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated remarkable efficacy in text embedding, yet current adaptation methods like LoRA face significant bottle

model-releasesarxiv-cs-ai
28 May 2026
Safety

ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation

DGX agent

arXiv:2605.28293v1 Announce Type: cross Abstract: Proactive Recommender Systems (PRSs) aim to guide user preference shift toward target items by generating paths of intermediate recommendations. Reinf

safetyarxiv-cs-ai
28 May 2026
Model Releases

ProvMind: Provenance-grounded reasoning for materials synthesis

DGX agent

arXiv:2605.28487v1 Announce Type: new Abstract: Materials process optimization requires reasoning over routes, conditions, tools and causal dependencies, yet most computational formulations flatten sy

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

PrunePath: Towards Highly Structured Sparse Language Models

DGX agent

arXiv:2605.28283v1 Announce Type: cross Abstract: Feed-forward networks (FFNs) dominate the parameter count and computation of modern language models, yet existing pruning methods often struggle to co

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Pruning and Distilling Mixture-of-Experts into Dense Language Models

DGX agent

arXiv:2605.28207v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory,

model-releasesarxiv-cs-ai
28 May 2026
Research

Quantum Machine Learning-based 6G edge Network: Enabling Adaptive Communication and Model Aggregation

DGX agent

arXiv:2605.27417v1 Announce Type: cross Abstract: With the advent of sixth-generation (6G) mobile communication technology, vehicle-to-everything (V2X) communication faces unprecedented challenges in

researcharxiv-cs-ai
28 May 2026
Applications

QuITE: Query-Based Irregular Time Series Embedding

DGX agent

arXiv:2605.28166v1 Announce Type: cross Abstract: Irregular Multivariate Time Series (IMTS) are common in practice, yet their irregular sampling complicates effective modeling. Existing approaches typ

applicationsarxiv-cs-ai
28 May 2026
Agents

RAG-Coding: Enhancing LLM Medical Coding with Structured External Knowledge

DGX agent

arXiv:2605.27377v1 Announce Type: cross Abstract: We present RAG-Coding, an agentic method for automated ICD-10-CM coding. RAG-Coding orchestrates four large language model (LLM) agents and grounds th

agentsarxiv-cs-ai
28 May 2026
Research

RAGe: A Retrieval-Augmented Generation Evaluation Framework

DGX agent

arXiv:2605.27445v1 Announce Type: cross Abstract: Deploying Large Language Model (LLM) applications, particularly those relying on Retrieval-Augmented Generation (RAG), remains challenging due to high

researcharxiv-cs-ai
28 May 2026
Safety

RE-TRIANGLE: Does TRIANGLE Enable Multimodal Alignment Beyond Cosine Similarity in Retrieval?

DGX agent

arXiv:2605.27436v1 Announce Type: cross Abstract: Multimodal alignment is critical for bridging the semantic gap in information retrieval. However, traditional pairwise strategies introduce a geometri

safetyarxiv-cs-ai
28 May 2026
Local Ai

Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions

DGX agent

arXiv:2605.27750v1 Announce Type: cross Abstract: Recent work has shown that Vision-Language Models (VLMs) used for optical character recognition (OCR) can generate plausible but visually unsupported

local-aiarxiv-cs-ai
28 May 2026
Agents

Reasoning and Planning with Dynamically Changing Norms

DGX agent

arXiv:2605.27622v1 Announce Type: new Abstract: To safely interact with humans, AI agents must both know our norms and consider them during planning. However, such norm-guided planning has been less e

agentsarxiv-cs-ai
28 May 2026
Safety

Reasoning Matters: Mitigate Hallucination in Multimodal Large Reasoning Models via Reasoning-Conditioned Preference Optimization

DGX agent

arXiv:2605.27906v1 Announce Type: new Abstract: Multimodal Large Reasoning Models introduce the reasoning paradigm, demonstrating strong capabilities on complex vision-language tasks. However, they st

safetyarxiv-cs-ai
28 May 2026
Applications

REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading

DGX agent

arXiv:2605.27402v1 Announce Type: cross Abstract: Open-ended grading is central to equitable and personalized education, yet manual grading remains time-consuming and costly, underscoring the need for

applicationsarxiv-cs-ai
28 May 2026
Model Releases

REED: Post-Training Representation Editing for Cross-Domain Linguistic Steganalysis

DGX agent

arXiv:2605.28298v1 Announce Type: new Abstract: In real-world scenarios of linguistic steganalysis, tested texts usually come from unseen domains with different vocabularies, topics, writing styles, a

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing

DGX agent

arXiv:2511.14584v3 Announce Type: replace-cross Abstract: We present ReflexGrad, a dual-process architecture for within-episode failure recovery in LLM agents without demonstrations. When agents commi

model-releasesarxiv-cs-ai
28 May 2026
Safety

Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

DGX agent

arXiv:2605.28553v1 Announce Type: new Abstract: In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on

safetyarxiv-cs-ai
28 May 2026
Model Releases

Regression Language Models for Code

DGX agent

arXiv:2509.26476v2 Announce Type: replace-cross Abstract: We study code-to-metric regression: predicting numeric outcomes of code executions, a challenging task due to the open-ended nature of program

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Relational Semantic Reasoning on 3D Scene Graphs for Open World Interactive Object Search

DGX agent

arXiv:2603.05642v2 Announce Type: replace-cross Abstract: Open-world interactive object search in household environments requires understanding semantic relationships between objects and their surroun

model-releasesarxiv-cs-ai
28 May 2026
Research

RelaxFlow: Text-Driven Amodal 3D Generation

DGX agent

arXiv:2603.05425v2 Announce Type: replace-cross Abstract: Image-to-3D generation faces inherent semantic ambiguity under occlusion, where partial observation alone is often insufficient to determine o

researcharxiv-cs-ai
28 May 2026
Model Releases

Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

DGX agent

arXiv:2605.28044v1 Announce Type: new Abstract: Cited RAG evaluation often treats visible sources as a grounding signal, but a real, topically relevant citation can still under-warrant the attached wo

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions

DGX agent

arXiv:2605.27819v1 Announce Type: cross Abstract: Sparse autoencoders are usually trained one layer at a time, even though transformer residual stream activations are strongly coupled across depth. Th

model-releasesarxiv-cs-ai
28 May 2026
Applications

ResearchLoop: An Evidence-Gated Control Plane for AI-Assisted Research

DGX agent

arXiv:2605.28282v1 Announce Type: new Abstract: AI-assisted research compresses ideation, implementation, evaluation, and manuscript writing into a single interactive loop. This compression is useful,

applicationsarxiv-cs-ai
28 May 2026
Research

Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models

DGX agent

arXiv:2605.27813v1 Announce Type: cross Abstract: Text-to-image diffusion models generate images through an iterative denoising process, so internal neural layers produce trajectories of activations r

researcharxiv-cs-ai
28 May 2026
Model Releases

Resource-Constrained Affect Modelling via Variance Regularisation Pruning

DGX agent

arXiv:2605.27479v1 Announce Type: cross Abstract: Affective computing systems are increasingly embedded in pervasive and interactive environments, such as adaptive games, assistive technologies, and r

model-releasesarxiv-cs-ai
28 May 2026
Safety

Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning

DGX agent

arXiv:2605.27765v1 Announce Type: cross Abstract: Self-Distillation Policy Optimization (SDPO) provides dense token-level credit assignment for reinforcement learning with large language models by lev

safetyarxiv-cs-ai
28 May 2026
Agents

Rethinking Memory as Continuously Evolving Connectivity

DGX agent

arXiv:2605.28773v1 Announce Type: cross Abstract: Existing memory-augmented LLM agents often treat memory as a static repository with pre-defined representations and fixed retrieval pipelines, which i

agentsarxiv-cs-ai
28 May 2026
Tutorials

Revealing Algorithmic Deductive Circuits for Logical Reasoning

DGX agent

arXiv:2605.27824v1 Announce Type: new Abstract: Recent studies have shown that Large Language Models (LLMs) can achieve strong reasoning performance by incorporating functional symbolic representation

tutorialsarxiv-cs-ai
28 May 2026
Local Ai

Reverse Probing: Supervised Token-level Uncertainty Quantification for Large Language Models in Clinical Text

DGX agent

arXiv:2605.28740v1 Announce Type: cross Abstract: As large language models are increasingly deployed for clinical text, ensuring they can reliably signal their own uncertainty becomes critical. Most e

local-aiarxiv-cs-ai
28 May 2026
Research

Revisiting Anthropomorphic Reflection Markers in Large Language Model Reasoning

DGX agent

arXiv:2605.28305v1 Announce Type: cross Abstract: Large Language Models (LLMs) often produce explicit reflective traces during complex reasoning, accompanied by anthropomorphic markers such as wait, h

researcharxiv-cs-ai
28 May 2026
Applications

Revisiting Change Detection Methods for their Application to Serac Fall Time-Lapse Monitoring

DGX agent

arXiv:2605.28100v1 Announce Type: cross Abstract: In an era where climate change aggravates environmental uncertainties, the identification and detection of event precursors are becoming crucial to mi

applicationsarxiv-cs-ai
28 May 2026
Research

Revisiting Graph Autoencoders as Implicit Contrastive Learners

DGX agent

arXiv:2410.10241v2 Announce Type: replace-cross Abstract: Graph autoencoders (GAEs) and graph contrastive learning (GCL) are two major paradigms for self-supervised representation learning on graphs,

researcharxiv-cs-ai
28 May 2026
← Previous
1…251252253254255…448
Next →