AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Safety

Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages

DGX agent

arXiv:2608.00533v1 Announce Type: new Abstract: Large Language Models have achieved substantial progress in reasoning capabilities. Yet in low-resource native settings, many suffer from cross-lingual

safetyarxiv-cs-cl
4 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Neural operator learning for collision-aware trajectory planning of spacecraft swarms

DGX agent

arXiv:2608.00320v1 Announce Type: new Abstract: Autonomous spacecraft swarms must plan fuel-efficient, collision-free maneuvers in increasingly congested orbits, yet classical trajectory optimization

safetyarxiv-cs-lg
4 Aug 2026
Research

PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters' Lack of Knowledge

DGX agent

arXiv:2608.01598v1 Announce Type: new Abstract: Simulating human-like Theory of Mind (ToM) has been a longstanding problem in natural language processing (NLP). To address this, existing works introdu

researcharxiv-cs-cl
4 Aug 2026
Model Releases

RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

DGX agent

arXiv:2608.01247v1 Announce Type: new Abstract: Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse und

model-releasesarxiv-cs-cl
4 Aug 2026
Safety

Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

DGX agent

arXiv:2608.01423v1 Announce Type: cross Abstract: Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing

safetyarxiv-cs-lg
4 Aug 2026
Local Ai

SVRepair: Structured Visual Reasoning for Automated Program Repair

DGX agent

arXiv:2602.06090v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently been applied to Automated Program Repair (APR), yet most existing approaches remain unimodal and fa

local-aiarxiv-cs-cv
4 Aug 2026
Model Releases

The Condition-Number Barrier in Sparse Least Squares

DGX agent

arXiv:2608.02588v1 Announce Type: cross Abstract: In [AS21], Axiotis and Sviridenko conjectured that the linear dependence on the restricted condition number in sparse convex optimization cannot be im

model-releasesarxiv-cs-lg
4 Aug 2026
Safety

Track-Guided Hierarchical Reinforcement Learning for Autonomous Vehicle Drifting with Minimum-Lap-Time Planning

DGX agent

arXiv:2608.00113v1 Announce Type: new Abstract: In Formula 1, drivers optimize racing lines within tire grip limits to minimize lap times; however, in rally racing, drivers intentionally break tractio

safetyarxiv-cs-ro
4 Aug 2026
Research

TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory

DGX agent

arXiv:2608.01922v1 Announce Type: new Abstract: Multimodal Large Reasoning Models (MLRMs) have achieved strong performance on tasks requiring visual understanding and multi-step inference. However, as

researcharxiv-cs-cl
4 Aug 2026
Model Releases

When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim Verification

DGX agent

arXiv:2608.01409v1 Announce Type: new Abstract: Biomedical fact-checking systems must do more than predict whether a claim is supported, contradicted, or unaddressed: they should also produce evidence

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Why Large Language Models Fail at Tabular Prediction

DGX agent

arXiv:2608.02412v1 Announce Type: new Abstract: Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the

model-releasesarxiv-cs-lg
4 Aug 2026
Safety

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

DGX agent

arXiv:2607.28890v1 Announce Type: cross Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes h

safetyarxiv-cs-ai
3 Aug 2026
Model Releases

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

DGX agent

arXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and compari

model-releasesarxiv-cs-ai
3 Aug 2026
Safety

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

DGX agent

arXiv:2607.28636v1 Announce Type: new Abstract: LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-drive

safetyarxiv-cs-cl
3 Aug 2026
Model Releases

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

DGX agent

arXiv:2607.29172v1 Announce Type: cross Abstract: While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-sourc

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Combining Large Language Models and Symbolic Reasoning for Multi-Robot Temporal Planning through Explainable Knowledge Bases

DGX agent

arXiv:2502.19135v2 Announce Type: replace Abstract: We present PLANTOR, a framework for generating and executing multi-robot task plans from natural-language task descriptions through LLM-assisted kno

model-releasesarxiv-cs-ai
3 Aug 2026
Local Ai

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

DGX agent

arXiv:2607.22186v2 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimiza

local-aiarxiv-cs-ai
3 Aug 2026
Research

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

DGX agent

arXiv:2607.29394v1 Announce Type: cross Abstract: Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based contrast ag

researcharxiv-cs-ai
3 Aug 2026
Safety

Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers

DGX agent

arXiv:2607.28904v1 Announce Type: cross Abstract: This paper presents a speculative Human-Computer Interaction design proposal for encouraging geopolitical reflexivity amongst tech workers at geopolit

safetyarxiv-cs-ai
3 Aug 2026
Safety

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

DGX agent

arXiv:2607.29246v1 Announce Type: new Abstract: Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a

safetyarxiv-cs-ai
3 Aug 2026
Research

Explore Beyond the Boundary Using Entropic Information

DGX agent

arXiv:2607.29419v1 Announce Type: cross Abstract: In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guid

researcharxiv-cs-ai
3 Aug 2026
Model Releases

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

DGX agent

arXiv:2607.16057v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as

model-releasesarxiv-cs-ai
3 Aug 2026
Safety

How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories

DGX agent

arXiv:2607.28674v1 Announce Type: new Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpr

safetyarxiv-cs-ai
3 Aug 2026
Model Releases

Predicting Steel Fatigue Life from Micrographs Using Physics-Informed Deep Learning

DGX agent

arXiv:2607.28695v1 Announce Type: cross Abstract: Here is the plain text version optimized for arXiv's submission form. Custom macros (like CV and SI) have been converted to standard text/math so they

model-releasesarxiv-cs-ai
3 Aug 2026
Research

Shall We Play a Game? Language Models for Open-ended Wargames

DGX agent

arXiv:2509.17192v3 Announce Type: replace Abstract: LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing

researcharxiv-cs-ai
3 Aug 2026
Model Releases

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models

DGX agent

arXiv:2603.06828v2 Announce Type: replace-cross Abstract: We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better. Standa

model-releasesarxiv-cs-ai
3 Aug 2026
Safety

TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking

DGX agent

arXiv:2607.28680v1 Announce Type: new Abstract: Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities. Existing approaches typically rely on

safetyarxiv-cs-cl
3 Aug 2026
Applications

The Checking Problem: What must be true before AI ships in a regulated firm

DGX agent

arXiv:2607.28666v1 Announce Type: new Abstract: Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. This paper measures the mechanism. Six document-heavy workflows of

applicationsarxiv-cs-cl
3 Aug 2026
Safety

Unified continuous-time q-learning for mean-field game and mean-field control problems

DGX agent

arXiv:2407.04521v3 Announce Type: replace-cross Abstract: This paper studies the continuous-time q-learning in mean-field jump-diffusion models in a setting where the environment simulator does not pr

safetyarxiv-cs-lg
3 Aug 2026
Model Releases

WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics

DGX agent

arXiv:2601.02430v3 Announce Type: replace-cross Abstract: Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and com

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations

DGX agent

arXiv:2607.28648v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on modifying cogn

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports

DGX agent

arXiv:2601.15297v3 Announce Type: replace Abstract: Reliable question answering over long institutional documents requires more than topical retrieval: a system must localize the exact passage that su

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

DGX agent

arXiv:2607.27845v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Continuous-time reinforcement learning for optimal switching over multiple regimes

DGX agent

arXiv:2512.04697v3 Announce Type: replace-cross Abstract: This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes. We consider a type

model-releasesarxiv-cs-lg
31 Jul 2026
Safety

Learning Social Robot Navigation By Sensing Human Legs

DGX agent

arXiv:2607.27922v1 Announce Type: new Abstract: Robots navigating among pedestrians typically sense their surroundings with a 2D LiDAR mounted close to the ground. At that height, the sensor mostly se

safetyarxiv-cs-ro
31 Jul 2026
Model Releases

MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

DGX agent

arXiv:2607.27895v1 Announce Type: cross Abstract: Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological s

model-releasesarxiv-cs-cv
31 Jul 2026
Safety

Optimizing Sensor Placement for Hydrogen Leak Detection in Enclosed Infrastructure: A Comparative Study Using CFD-informed Genetic Algorithm and DeepSets Neural Surrogate

DGX agent

arXiv:2607.26078v1 Announce Type: cross Abstract: Hydrogen infrastructure in enclosed environments, such as parking facilities for fuel cell vehicles, presents significant safety challenges due to hyd

safetyarxiv-cs-ai
31 Jul 2026
Model Releases

PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform's Perspective

DGX agent

arXiv:2607.27265v1 Announce Type: new Abstract: Real-time bidding is central to computational advertising, comprising three elements: Supply Side Platform (SSP) selling ad impressions, Demand Side Pla

model-releasesarxiv-cs-lg
31 Jul 2026
Safety

Policy Gradient Steering: Interventions from Behavioral Objectives

DGX agent

arXiv:2607.27574v1 Announce Type: new Abstract: Activation steering has emerged in large language models as a lightweight alternative for dynamically changing a model's behavior at inference time. How

safetyarxiv-cs-lg
31 Jul 2026
Applications

Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation

DGX agent

arXiv:2607.27210v1 Announce Type: new Abstract: The exponential growth of scholarly publications requires automated tools for effective information synthesis. However, simple, single-shot prompting me

applicationsarxiv-cs-cl
31 Jul 2026
Model Releases

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

DGX agent

arXiv:2607.28509v1 Announce Type: new Abstract: Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference

model-releasesarxiv-cs-cv
31 Jul 2026
Tutorials

Relational Scene Graphs for Object Grounding of Natural Language Commands

DGX agent

arXiv:2602.04635v2 Announce Type: replace Abstract: Robots are finding wider adoption in human environments, increasing the need for natural human-robot interaction. However, understanding a natural l

tutorialsarxiv-cs-ro
31 Jul 2026
Model Releases

Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models

DGX agent

arXiv:2607.27384v1 Announce Type: new Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

STEREODISCO: Discovering Stereotypicality in LLMs

DGX agent

arXiv:2607.27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psycholog

model-releasesarxiv-cs-lg
31 Jul 2026
Safety

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

DGX agent

arXiv:2607.28590v1 Announce Type: cross Abstract: Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view t

safetyarxiv-cs-cl
31 Jul 2026
Safety

BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

DGX agent

arXiv:2607.26914v1 Announce Type: new Abstract: Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed

safetyarxiv-cs-ro
30 Jul 2026
Local Ai

Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations

DGX agent

arXiv:2607.26481v1 Announce Type: new Abstract: Detecting when the statistical behavior of an engineered system changes, and identifying which component is responsible, are core problems in the monito

local-aiarxiv-cs-lg
30 Jul 2026
Safety

Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making

DGX agent

arXiv:2607.27022v1 Announce Type: new Abstract: Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet

safetyarxiv-cs-cl
30 Jul 2026
← Previous
1…205206207208209…233
Next →