AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
22,134 results
Safety

A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach

DGX agent

arXiv:2512.11944v2 Announce Type: replace-cross Abstract: Motion planning for autonomous driving (AD) faces a critical trade-off. While traditional rule-based pipelines offer verifiable safety and int

safetyarxiv-cs-ai
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

AgentSchool: An LLM-Powered Multi-Agent Simulation for Education

DGX agent

arXiv:2605.30144v1 Announce Type: new Abstract: Despite the rapid deployment of LLMs into classrooms, validating educational AI remains uniquely intractable: interventions act on developing learners w

agentsarxiv-cs-ai
29 May 2026
Agents

An Approach for Thyroid Nodule Analysis Using Thermographic Images

DGX agent

arXiv:2605.29221v1 Announce Type: new Abstract: Thyroid cancer is said to be the second most common type of cancer in female individuals and the third in males by 2030, according to projections. In ge

agentsarxiv-cs-cv
29 May 2026
Model Releases

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

DGX agent

arXiv:2605.29462v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified i

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

DGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

DGX agent

arXiv:2605.30188v1 Announce Type: cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibr

model-releasesarxiv-cs-ai
29 May 2026
Agents

Catalyst-Agent: Autonomous heterogeneous catalyst screening with an LLM Agent

DGX agent

arXiv:2603.01311v2 Announce Type: replace Abstract: The discovery of novel catalysts tailored for particular applications is a major challenge for the twenty-first century. Traditional methods for thi

agentsarxiv-cs-cl
29 May 2026
Model Releases

Comparative Evaluation of Machine Translation Systems on Images with Text

DGX agent

arXiv:2605.29476v1 Announce Type: new Abstract: This work presents a comparative evaluation of machine translation systems applied to images containing textual information, a task that lies at the int

model-releasesarxiv-cs-cl
29 May 2026
Agents

CONCAT: Consensus- and Confidence-Driven Ad Hoc Teaming for Efficient LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.29612v1 Announce Type: cross Abstract: Although large language model (LLM) based multi-agent systems (MAS) show their capability to solve complex tasks and achieve higher performance over s

agentsarxiv-cs-cl
29 May 2026
Model Releases

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective

DGX agent

arXiv:2502.03805v2 Announce Type: replace Abstract: Large language models have revolutionized natural language processing but face significant challenges of high storage and runtime costs, due to the

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking

DGX agent

arXiv:2605.30107v1 Announce Type: new Abstract: Creating spoken dialogue datasets is methodologically challenging, and these challenges are amplified when the goal is to build multilingual, multi-para

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning

DGX agent

arXiv:2605.29339v1 Announce Type: new Abstract: With the rapid advancement of multimodal large language models (MLLMs), models have demonstrated increasingly powerful multimodal capabilities. However,

model-releasesarxiv-cs-cv
29 May 2026
Safety

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

DGX agent

arXiv:2510.27607v3 Announce Type: replace Abstract: Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predictin

safetyarxiv-cs-cv
29 May 2026
Model Releases

DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

DGX agent

arXiv:2605.29256v1 Announce Type: cross Abstract: Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality

model-releasesarxiv-cs-ai
29 May 2026
Agents

E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing

DGX agent

arXiv:2512.03109v2 Announce Type: replace-cross Abstract: Agentic AI systems execute a sequence of actions, such as reasoning steps or tool calls, in response to a user prompt. To evaluate the success

agentsarxiv-cs-ai
29 May 2026
Model Releases

EarthShift: a benchmark for measuring robustness to real-world distribution shifts in Earth observation

DGX agent

arXiv:2605.29330v1 Announce Type: new Abstract: Current Earth observation benchmarks focus on measuring performance on diverse tasks and applications, typically measuring generalization in-distributio

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Evaluating Dataset Watermarking for Fine-tuning Traceability of Customized Diffusion Models: A Comprehensive Benchmark and Removal Approach

DGX agent

arXiv:2511.19316v2 Announce Type: replace-cross Abstract: Recent fine-tuning techniques for diffusion models enable them to reproduce specific image sets, such as particular faces or artistic styles,

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Evolutionary Rule Extraction from Corporate Default Prediction Models

DGX agent

arXiv:2605.29478v1 Announce Type: cross Abstract: Small and medium-sized enterprises (SMEs) represent the majority of firms in most economies and often face financial constraints and higher vulnerabil

model-releasesarxiv-cs-ai
29 May 2026
Local Ai

FHRFormer: A Self-Supervised Masked Transformer Framework for Fetal Heart Rate Time-Series Inpainting and Forecasting

DGX agent

arXiv:2605.29695v1 Announce Type: new Abstract: Approximately 10% of newborns require assistance to initiate breathing at birth, and around 5% need ventilation support. Fetal heart rate (FHR) monitori

local-aiarxiv-cs-ai
29 May 2026
Agents

Formalizing Mathematics at Scale

DGX agent

arXiv:2605.29955v1 Announce Type: new Abstract: We present AutoformBot, a multi-agent system for building an Autoformalized Textbook Library At Scale (Atlas) in Lean 4. AutoformBot orchestrates thousa

agentsarxiv-cs-ai
29 May 2026
Model Releases

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

DGX agent

arXiv:2605.29107v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a gr

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GPIC: A Giant Permissive Image Corpus for Visual Generation

DGX agent

arXiv:2605.30341v1 Announce Type: cross Abstract: Studying scalable methods for visual generative modeling requires large, accessible, and stable datasets. We introduce GPIC, a Giant Permissive Image

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Gram: Assessing sabotage propensities via automated alignment auditing

DGX agent

arXiv:2605.30322v1 Announce Type: cross Abstract: We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models ac

model-releasesarxiv-cs-ai
29 May 2026
Safety

GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German

DGX agent

arXiv:2605.30214v1 Announce Type: new Abstract: Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about referenc

safetyarxiv-cs-cl
29 May 2026
Agents

Human-in-the-Loop Swarms: A Bionic Swarm Approach to Real-World Soil Mapping

DGX agent

arXiv:2605.29091v1 Announce Type: new Abstract: Swarm and field robotics face significant barriers to real-world validation due to the high cost and development time to deploy hardware. This paper int

agentsarxiv-cs-ro
29 May 2026
Model Releases

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

DGX agent

arXiv:2605.29473v1 Announce Type: cross Abstract: Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond inf

model-releasesarxiv-cs-ai
29 May 2026
Applications

LsrIF: Enhancing Logic-Structured Instruction Following of Large Language Models

DGX agent

arXiv:2601.06431v3 Announce Type: replace Abstract: Instruction following is critical for large language models, yet real-world instructions often involve multiple constraints with logical structures,

applicationsarxiv-cs-ai
29 May 2026
Model Releases

MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

DGX agent

arXiv:2601.04633v2 Announce Type: replace Abstract: Machine-Generated Text (MGT) is becoming increasingly difficult to distinguish from Human-Written Text (HWT). This trend has exacerbated malicious a

model-releasesarxiv-cs-cl
29 May 2026
Tutorials

MonoDuo: Using One Robot Arm to Learn Bimanual Policies

DGX agent

arXiv:2605.29298v1 Announce Type: new Abstract: Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual r

tutorialsarxiv-cs-ro
29 May 2026
Agents

MOOSE-Copilot: A Web-Based Interactive Assistant for Unified Exploratory and Fine-Grained Scientific Hypothesis Discovery

DGX agent

arXiv:2605.29475v1 Announce Type: cross Abstract: Large language models (LLMs) show remarkable potential in scientific hypothesis discovery. However, existing approaches face two critical limitations:

agentsarxiv-cs-ai
29 May 2026
Model Releases

Multimodal LLMs See Sentiment

DGX agent

arXiv:2508.16873v3 Announce Type: replace Abstract: Understanding how visual content conveys sentiment is increasingly important in a digital landscape dominated by imagery. However, sentiment percept

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

DGX agent

arXiv:2605.29300v1 Announce Type: cross Abstract: Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses ar

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Personalized Turn-Level User Conversation Satisfaction Benchmark

DGX agent

arXiv:2605.29711v1 Announce Type: cross Abstract: User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user

model-releasesarxiv-cs-ai
29 May 2026
Safety

Practitioner Beliefs and Behaviors in AI-Enhanced Education: DOT Framework Survey Evidence

DGX agent

arXiv:2605.29041v1 Announce Type: new Abstract: This study reports findings from a cross-sectional survey (n = 72) of higher education practitioners examining beliefs, behaviors, and institutional con

safetyarxiv-cs-ai
29 May 2026
Safety

Provably Secure Agent Guardrail

DGX agent

arXiv:2605.29251v1 Announce Type: new Abstract: As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates

safetyarxiv-cs-ai
29 May 2026
Model Releases

PTCG-Bench: Can LLM Agents Master Pokemon Trading Card Game?

DGX agent

arXiv:2605.29653v1 Announce Type: new Abstract: Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds. Autonomous agents require sim

model-releasesarxiv-cs-ai
29 May 2026
Safety

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

DGX agent

arXiv:2605.28899v1 Announce Type: cross Abstract: Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses si

safetyarxiv-cs-ai
29 May 2026
Model Releases

RAISE: RAG Design as an Architecture Search Problem

DGX agent

arXiv:2605.30029v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems expose numerous design choices spanning query rewriting, chunking, retrieval depth, reranking, and context

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Realistic honeypot evaluations for scheming propensity

DGX agent

arXiv:2605.29729v1 Announce Type: new Abstract: We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming

model-releasesarxiv-cs-lg
29 May 2026
Safety

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

DGX agent

arXiv:2602.20141v2 Announce Type: replace Abstract: Mean Field Games (MFGs) provide a principled framework for modelling interactions in large population systems. However, algorithmic progress has bee

safetyarxiv-cs-ai
29 May 2026
Safety

Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents

DGX agent

arXiv:2605.28850v1 Announce Type: new Abstract: We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. Using TradeArena, an

safetyarxiv-cs-lg
29 May 2026
Model Releases

TAE: Target-aware enhancer for nighttime UAV tracking

DGX agent

arXiv:2605.29558v1 Announce Type: new Abstract: Severe image degradation under low-light nighttime conditions constitutes a core bottleneck preventing all-day applications for UAV-based single object

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Temporal Stability and Few-Shot Prompting in Math Task Assessment

DGX agent

arXiv:2605.30151v1 Announce Type: new Abstract: As AI tools become increasingly integrated into educational contexts, questions arise about both their stability over time and their responsiveness to p

model-releasesarxiv-cs-ai
29 May 2026
Agents

AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models

DGX agent

arXiv:2605.27873v1 Announce Type: new Abstract: AI models underpin data-centric applications from image and text processing to scientific discovery in biology, physics, and chemistry. Yet developing t

agentsarxiv-cs-ai
28 May 2026
Model Releases

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

DGX agent

arXiv:2605.28642v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment par

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Benchmarking Fairness in Spiking Neural Networks: Data Bias, Spurious Features, and Hardware Effects

DGX agent

arXiv:2605.27407v1 Announce Type: cross Abstract: Evaluating fairness in Spiking Neural Networks (SNNs) demands rigorous benchmarks that reflect real-world complexities, yet existing assessments remai

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Category-Level 3D Correspondence in Camera Space via Morphable Object Priors

DGX agent

arXiv:2605.28257v1 Announce Type: new Abstract: Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estim

model-releasesarxiv-cs-cv
28 May 2026
Safety

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

DGX agent

arXiv:2603.22335v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior dis

safetyarxiv-cs-ai
28 May 2026
← Previous
1…431432433434435…462
Next →