AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,189 results
Research

Are Researchers Being Replaced by Artificial Intelligence?

DGX agent

arXiv:2605.16294v1 Announce Type: cross Abstract: A Nature survey from 2023 involving 1,600 researchers shows that scientists are ``concerned, as well as excited, by the increasing use of artificial-i

researcharxiv-cs-ai
19 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents

DGX agent

arXiv:2605.17324v1 Announce Type: cross Abstract: Clarification-seeking behavior is widely regarded as a desirable property of LLM agents, enabling them to resolve ambiguity before acting on underspec

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting

DGX agent

arXiv:2605.17937v1 Announce Type: cross Abstract: Quantitative backtesting is essential for evaluating trading strategies but remains hampered by high technical barriers and limited scalability. While

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Benchmarking Mythos-Linked Bug Rediscovery

DGX agent

arXiv:2605.17416v1 Announce Type: cross Abstract: Anthropic's April 2026 Mythos materials combine benchmark claims with concrete bug-finding stories across OpenBSD, FreeBSD, Linux, FFmpeg, and browser

model-releasesarxiv-cs-ai
19 May 2026
Research

Beyond Morphology: Quantifying the Diagnostic Power of Color Features in Cancer Classification

DGX agent

arXiv:2605.18522v1 Announce Type: cross Abstract: In histopathology, human experts primarily rely on color as a means of enhancing contrast to interpret tissue morphology, whereas machine vision model

researcharxiv-cs-ai
19 May 2026
Research

Bridging the Gap: Converting Read Text to Conversational Dialogue

DGX agent

arXiv:2605.18001v1 Announce Type: new Abstract: In recent advancements within speech processing, converting read speech to conversational speech has gained significant attention. The primary challenge

researcharxiv-cs-cl
19 May 2026
Safety

Building Reliable Arithmetic Multipliers Under NBTI Aging and Process Variations

DGX agent

arXiv:2605.18444v1 Announce Type: cross Abstract: Hardware aging poses a significant challenge for integrated circuits (ICs), leading to performance degradation and eventual failure. In this work, we

safetyarxiv-cs-ai
19 May 2026
Model Releases

Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows

DGX agent

arXiv:2605.18327v1 Announce Type: new Abstract: AI agents deployed into SRE workflows currently derive their understanding of environment state from raw observability telemetry at query time, paying a

model-releasesarxiv-cs-ai
19 May 2026
Safety

Code as Agent Harness

DGX agent

arXiv:2605.18747v1 Announce Type: cross Abstract: Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to reposi

safetyarxiv-cs-ai
19 May 2026
Safety

Convex Dataset Valuation for Post-Training

DGX agent

arXiv:2605.16704v1 Announce Type: new Abstract: Improving LLM performance on downstream tasks sometimes requires leveraging auxiliary datasets during post-training. In practice, however, developers fa

safetyarxiv-cs-lg
19 May 2026
Model Releases

CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability

DGX agent

arXiv:2602.03012v2 Announce Type: replace-cross Abstract: Evaluating and improving the security capabilities of code agents requires high-quality, executable vulnerability tasks. However, existing wor

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs

DGX agent

arXiv:2605.18498v1 Announce Type: cross Abstract: Expert specialization in Mixture-of-Experts (MoE) models remains poorly understood, with traditional evaluations conflating architectural load-balanci

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

EndoCogniAgent: Closed-Loop Agentic Reasoning with Self-Consistency Validation for Endoscopic Diagnosis

DGX agent

arXiv:2508.07292v3 Announce Type: replace Abstract: Endoscopic diagnosis is an iterative process in which clinicians progressively acquire, compare, and verify local visual evidence before reaching a

model-releasesarxiv-cs-ai
19 May 2026
Safety

Estimating Item Difficulty with Large Language Models as Experts

DGX agent

arXiv:2605.18562v1 Announce Type: cross Abstract: Accurate estimates of item difficulty are essential for valid assessment and effective adaptive learning. However, for newly created tasks, response d

safetyarxiv-cs-ai
19 May 2026
Model Releases

Evaluating Cognitive Age Alignment in Interactive AI Agents

DGX agent

arXiv:2605.17894v1 Announce Type: new Abstract: While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across doma

model-releasesarxiv-cs-ai
19 May 2026
Safety

Event-Grounded Sparse Autoencoders for Vision-Language-Action Policies

DGX agent

arXiv:2605.17204v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies translate language and visual inputs into robot actions, where their hidden representations directly shape close

safetyarxiv-cs-ai
19 May 2026
Research

Generalized Functional ANOVA in Closed-Form: A Unified View of Additive Explanations

DGX agent

arXiv:2605.18422v1 Announce Type: cross Abstract: The functional ANOVA, or Hoeffding decomposition, provides a principled framework for interpretability by decomposing a model prediction into main eff

researcharxiv-cs-lg
19 May 2026
Model Releases

Generative Artificial Intelligence for Literature Reviews

DGX agent

arXiv:2605.16475v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI), based on large-language models (LLMs), such as ChatGPT, has taken organizations, academia, and the public

model-releasesarxiv-cs-cl
19 May 2026
Agents

Geometry-Aware Surrogate for Real-Time Hydrodynamics Estimation of Autonomous Ground Vehicles in Amphibious Environments

DGX agent

arXiv:2605.18543v1 Announce Type: new Abstract: Autonomous ground vehicles operating in shallow water or flood-prone terrains require dynamic models that account for hydrodynamic forces. However, the

agentsarxiv-cs-ro
19 May 2026
Tutorials

GUIDE-VAE: Advancing Data Generation with User Information and Pattern Dictionaries

DGX agent

arXiv:2411.03936v2 Announce Type: replace Abstract: Generative modelling of multi-user datasets has become prominent in science and engineering. Generating a data point for a given user requires emplo

tutorialsarxiv-cs-lg
19 May 2026
Agents

Heterogeneous Information-Bottleneck Coordination Graphs for Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.17393v1 Announce Type: new Abstract: Coordination graphs are a central abstraction in cooperative multi-agent reinforcement learning (MARL), yet existing sparse-graph learners lack a theore

agentsarxiv-cs-ai
19 May 2026
Model Releases

HTSC-2025: A Benchmark Dataset of Ambient-Pressure High-Temperature Superconductors for AI-Driven Critical Temperature Prediction

DGX agent

arXiv:2506.03837v2 Announce Type: replace-cross Abstract: The discovery of high-temperature superconducting materials holds great significance for human industry and daily life. In recent years, resea

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools

DGX agent

arXiv:2605.17106v1 Announce Type: new Abstract: Production LLM deployments increasingly maintain heterogeneous model pools spanning order-of-magnitude cost differences. Existing routers make binary st

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

iMiGUE-3K: A Large-Scale Benchmark for Micro-Gesture Analysis with Self-Supervised Learning

DGX agent

arXiv:2605.17179v1 Announce Type: new Abstract: Emotion understanding is a fundamental challenge in affective computing and artificial intelligence. While existing approaches predominantly focus on fa

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025

DGX agent

arXiv:2305.07152v4 Announce Type: replace Abstract: Robotic assisted (RA) surgery promises to transform surgical intervention. Intuitive Surgical is committed to fostering these changes and the machin

model-releasesarxiv-cs-cv
19 May 2026
Tutorials

Inventorship in AI-Assisted Inventions: Designing an Experiment to Shape Case Law

DGX agent

arXiv:2605.16528v1 Announce Type: cross Abstract: The latest improvements in artificial intelligence (AI) raise new challenges for intellectual property laws, particularly concerning the inventorship

tutorialsarxiv-cs-ai
19 May 2026
Agents

LARGER: Lexically Anchored Repository Graph Exploration and Retrieval

DGX agent

arXiv:2605.16352v1 Announce Type: cross Abstract: Repository-level coding agents must first localize the files and symbols relevant to a task; failures at this stage can cascade across downstream obje

agentsarxiv-cs-ai
19 May 2026
Model Releases

LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning

DGX agent

arXiv:2605.16675v1 Announce Type: new Abstract: We introduce LinAlg-Bench, a diagnostic benchmark evaluating 10 frontier large language models on structured linear algebra computation across a strict

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

LiTS: A Modular Framework for LLM Tree Search

DGX agent

arXiv:2603.00631v2 Announce Type: replace Abstract: LiTS is a modular Python framework for LLM reasoning via tree search. It decomposes tree search into three reusable components (Policy, Transition,

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

LLMs in Qualitative Research: Opportunities, Limitations, and Practical Considerations

DGX agent

arXiv:2605.16538v1 Announce Type: cross Abstract: This paper examines the opportunities, limitations, and practical considerations associated with the use of large language models (LLMs) in qualitativ

model-releasesarxiv-cs-cl
19 May 2026
Safety

Measuring Changes in Instructor Class Design and Student Learning After the Release of Large Language Models (LLMs)

DGX agent

arXiv:2605.16284v1 Announce Type: cross Abstract: Student use of Generative AI (GenAI) products in completing their classwork, with or without their professors' knowledge and/or approval, has resulted

safetyarxiv-cs-ai
19 May 2026
Research

Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex

DGX agent

arXiv:2605.16468v1 Announce Type: cross Abstract: A central goal in understanding human vision is to uncover the visual features that drive neuronal activity. A growing body of work has used artificia

researcharxiv-cs-ai
19 May 2026
Agents

MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repair

DGX agent

arXiv:2605.17444v1 Announce Type: cross Abstract: Modern software ecosystems face a rapidly growing number of disclosed vulnerabilities, increasing the need for automated repair techniques that can op

agentsarxiv-cs-ai
19 May 2026
Safety

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

DGX agent

arXiv:2605.18549v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not alwa

safetyarxiv-cs-cl
19 May 2026
Safety

MSIQ: Moment-based Scale-Invariant Quality Measure for Single Image Super-Resolution

DGX agent

arXiv:2605.17588v1 Announce Type: new Abstract: Assessing the quality of single image super-resolution (SISR) results remains an open methodological problem. Common full-reference metrics (PSNR, SSIM,

safetyarxiv-cs-cv
19 May 2026
Safety

NEWTON: Agentic Planning for Physically Grounded Video Generation

DGX agent

arXiv:2605.18396v1 Announce Type: new Abstract: Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves on

safetyarxiv-cs-cv
19 May 2026
Safety

oldsymbol{f}-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control

DGX agent

arXiv:2605.17862v1 Announce Type: cross Abstract: Scaling on-policy distillation (OPD) for large language models (LLMs) confronts a fundamental tension: asynchronous execution is necessary for system

safetyarxiv-cs-ai
19 May 2026
Safety

OrbiSim: World Models as Differentiable Physics Engines for Embodied Intelligence

DGX agent

arXiv:2605.16395v1 Announce Type: cross Abstract: We present OrbiSim, a novel robotic simulation paradigm that redefines world models as a fully differentiable physics engine for embodied intelligence

safetyarxiv-cs-lg
19 May 2026
Model Releases

Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks

DGX agent

arXiv:2605.18583v1 Announce Type: cross Abstract: Coding agents now run autonomously with shell, file, and network privileges. When a user issues a benign request, the agent sometimes does more than a

model-releasesarxiv-cs-ai
19 May 2026
Safety

Parameterized 4-Qubit EWL Quantum Game Circuits with Dirac-Solow-Swan Hamiltonian Integration for Quadruple Helix Disruptive Innovation Recommender Systems

DGX agent

arXiv:2605.18080v1 Announce Type: cross Abstract: We present a novel parameterized 4-qubit Eisert-Wilkens-Lewenstein (EWL) quantum game circuit for recommender systems in quadruple helix innovation ec

safetyarxiv-cs-ai
19 May 2026
Agents

PULSE: Agentic Investigation with Passive Sensing for Proactive Intervention in Cancer Survivorship

DGX agent

arXiv:2605.17679v1 Announce Type: cross Abstract: Cancer survivors face elevated rates of depression, anxiety, and general emotional distress, yet the precise moments they most need support are often

agentsarxiv-cs-ai
19 May 2026
Local Ai

R2V Agent: Teaching SLMs When to Ask for Help

DGX agent

arXiv:2605.16604v1 Announce Type: new Abstract: Efficient agentic systems should incur expensive frontier-model costs only on decisions where a cheaper local model is likely to fail. Existing LLM casc

local-aiarxiv-cs-lg
19 May 2026
Agents

RAGA: Reading-And-Graph-building-Agent for Autonomous Knowledge Graph Construction and Retrieval-Augmented Generation

DGX agent

arXiv:2605.17072v1 Announce Type: new Abstract: Existing LLM-driven knowledge graph (KG) construction methods predominantly employ stateless batch processing pipelines, exhibiting structural deficienc

agentsarxiv-cs-ai
19 May 2026
Model Releases

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts

DGX agent

arXiv:2510.07239v2 Announce Type: replace Abstract: Automated red-teaming has emerged as a scalable approach for auditing Large Language Models (LLMs) prior to deployment, yet existing approaches lack

model-releasesarxiv-cs-cl
19 May 2026
Tutorials

Rover: Context-aware Conflict Resolution with LLM

DGX agent

arXiv:2605.17279v1 Announce Type: cross Abstract: Code merging is a significant challenge, particularly in large-scale projects. Existing solutions, including program analysis and machine learning, sh

tutorialsarxiv-cs-ai
19 May 2026
Agents

Same Signal, Different Semantics: A Cross-Framework Behavioral Analysis of Software Engineering Agents

DGX agent

arXiv:2605.18332v1 Announce Type: cross Abstract: Behavioral studies of LLM-based software engineering agents extract operational rules about which trajectory shapes correlate with higher resolution r

agentsarxiv-cs-ai
19 May 2026
Model Releases

Scalable Knowledge Editing for Mixture-of-Experts LLMs via Tensor-Structured Updates

DGX agent

arXiv:2605.16686v1 Announce Type: new Abstract: Knowledge editing (KE) provides a lightweight alternative to repeated fine-tuning of LLMs. However, most existing KE methods target dense feed-forward l

model-releasesarxiv-cs-lg
19 May 2026
Safety

Scale Determines Whether Language Models Organize Representation Geometry for Prediction

DGX agent

arXiv:2605.17084v1 Announce Type: cross Abstract: In language models, what a representation encodes is determined by the geometry of its representation space: distances, not activations, carry meaning

safetyarxiv-cs-cl
19 May 2026
← Previous
1…8687888990…109
Next →