AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
21,937 results
Applications

Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries

DGX agent

arXiv:2608.09532v1 Announce Type: cross Abstract: Enterprises increasingly seek to query data lakes using natural language via AI-driven tools like semantic operators or deep research agents. However,

applicationsarxiv-cs-ai
11 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

Explainable Machine Learning in Healthcare: Methods, Interpretation, and Applications for Clinical Research

DGX agent

arXiv:2608.07522v1 Announce Type: cross Abstract: We present a structured review of commonly used Explainable machine learning (XML) methodologies, including global and local interpretability tools su

local-aiarxiv-cs-lg
11 Aug 2026
Model Releases

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

DGX agent

arXiv:2608.09538v1 Announce Type: cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Enhancing Anomaly Resilience in Research Networks: A Large-Scale Forecasting Benchmark for Dynamic Security Baselining

DGX agent

arXiv:2608.05605v1 Announce Type: cross Abstract: Research and Education Networks (RENs) serve as critical infrastructure for scientific discovery, yet they face a unique security paradox: their norma

model-releasesarxiv-cs-lg
7 Aug 2026
Agents

Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents

DGX agent

arXiv:2608.02751v1 Announce Type: cross Abstract: Existing deep-research agents use a search-visit workflow that retrieves and reads whole pages, without considering the addressable structure that web

agentsarxiv-cs-ai
5 Aug 2026
Local Ai

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models

DGX agent

arXiv:2608.00019v1 Announce Type: new Abstract: Deploying large language models (LLMs) for operations research (OR) tasks remains challenging because correctness depends on a coherent modeling process

local-aiarxiv-cs-lg
4 Aug 2026
Model Releases

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

DGX agent

arXiv:2607.24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis

DGX agent

arXiv:2607.13602v1 Announce Type: cross Abstract: Systematic comparisons between current situations and structurally similar past events in the historical, i.e., historical analogies, is among the mos

model-releasesarxiv-cs-lg
16 Jul 2026
Safety

DR-Arena: an Automated Evaluation Framework for Deep Research Agents

DGX agent

arXiv:2601.10504v2 Announce Type: replace Abstract: As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, rel

safetyarxiv-cs-cl
10 Jul 2026
Model Releases

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

DGX agent

arXiv:2509.26076v2 Announce Type: replace Abstract: As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on researc

model-releasesarxiv-cs-cl
10 Jul 2026
Safety

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

DGX agent

arXiv:2607.07663v1 Announce Type: new Abstract: AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data t

safetyarxiv-cs-ai
9 Jul 2026
Model Releases

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

DGX agent

arXiv:2607.04438v1 Announce Type: cross Abstract: Research dissemination, turning a paper into a poster, a talk video, and a blog post, is still a manual last mile. Prior automation treats each artifa

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning

DGX agent

arXiv:2607.02927v1 Announce Type: cross Abstract: Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (V

model-releasesarxiv-cs-ai
7 Jul 2026
Agents

DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification

DGX agent

arXiv:2606.29746v1 Announce Type: new Abstract: Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge bases, remains

agentsarxiv-cs-ai
30 Jun 2026
Model Releases

Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop

DGX agent

arXiv:2606.29717v1 Announce Type: cross Abstract: Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of work has pr

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

COOPA: A Modular LLM Agent Architecture for Operations Research Problems

DGX agent

arXiv:2606.27611v1 Announce Type: new Abstract: Operations Research (OR) provides a rigorous framework for high-stakes decision-making, but effective OR modeling requires substantial domain knowledge,

agentsarxiv-cs-lg
29 Jun 2026
Model Releases

ReportLogic: Evaluating Logical Quality in Deep Research Reports

DGX agent

arXiv:2602.18446v2 Announce Type: replace-cross Abstract: Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports th

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Failure Modes of Large Language Models on Research-Level Mathematics: A Taxonomy and an Empirical Characterisation

DGX agent

arXiv:2606.24902v1 Announce Type: cross Abstract: The 'First Proof' benchmark [1] posed ten research-level mathematics questions to the strongest publicly available LLMs and found them consistently wr

model-releasesarxiv-cs-ai
25 Jun 2026
Safety

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?

DGX agent

arXiv:2510.02660v2 Announce Type: replace-cross Abstract: When researchers claim AI systems possess ToM or mental models, they are fundamentally discussing behavioral predictions and bias corrections

safetyarxiv-cs-ai
11 Jun 2026
Tutorials

Aesthetic Perspectives in Information Systems Research: A Hermeneutic Analysis

DGX agent

arXiv:2606.09839v1 Announce Type: cross Abstract: How might implicit aesthetic perspectives shape what Information Systems (IS) scholarship recognises as worthy of study (or not)? In this hermeneutic

tutorialsarxiv-cs-ai
10 Jun 2026
Safety

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence

DGX agent

arXiv:2606.07612v1 Announce Type: cross Abstract: We argue that many Anthropomorphic Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Automatic Generation of Titles for Research Papers Using Language Models

DGX agent

arXiv:2606.05085v1 Announce Type: cross Abstract: The title of a research paper conveys its primary idea and, occasionally, its conclusions in a clear and concise manner. Choosing an appropriate title

model-releasesarxiv-cs-ai
4 Jun 2026
Agents

MIRROR: A Multi-Agent Framework with Iterative Adaptive Revision and Hierarchical Retrieval for Optimization Modeling in Operations Research

DGX agent

arXiv:2602.03318v3 Announce Type: replace Abstract: Operations Research (OR) relies on expert-driven modeling-a slow and fragile process ill-suited to novel scenarios. While large language models (LLM

agentsarxiv-cs-cl
2 Jun 2026
Model Releases

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

DGX agent

arXiv:2605.29468v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (

model-releasesarxiv-cs-ai
29 May 2026
Agents

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

DGX agent

arXiv:2605.28003v1 Announce Type: new Abstract: The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully en

agentsarxiv-cs-cl
28 May 2026
Safety

Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification

DGX agent

arXiv:2510.00902v2 Announce Type: replace Abstract: Transfer learning is crucial for medical imaging, yet the selection of source datasets often relies on researchers' intuition rather than systematic

safetyarxiv-cs-cv
27 May 2026
Model Releases

ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research

DGX agent

arXiv:2601.21008v3 Announce Type: replace-cross Abstract: Operations Research practitioners debug infeasible models through an iterative process: inspecting Irreducible Infeasible Subsystems ( IIS), i

model-releasesarxiv-cs-ai
27 May 2026
Local Ai

Persistent AI Agents in Academic Research: A Single-Investigator Implementation Case Study

DGX agent

arXiv:2605.26870v1 Announce Type: cross Abstract: Background: Large language models are typically evaluated as models, benchmarks, or short conversational episodes. Less is known about what happens wh

local-aiarxiv-cs-ai
27 May 2026
Model Releases

VeriTrace: Evolving Mental Models for Deep Research Agents

DGX agent

arXiv:2605.26081v1 Announce Type: new Abstract: Deep research agents face vast, interdependent, and pervasively uncertain information. Existing systems explore what evolving intermediate representatio

model-releasesarxiv-cs-ai
26 May 2026
Safety

Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses

DGX agent

arXiv:2605.19229v1 Announce Type: new Abstract: Survey research faces mounting structural challenges: declining response rates, sample bias, block-wise missingness among at-risk respondents, and AI-as

safetyarxiv-cs-ai
20 May 2026
Model Releases

LLMs in Qualitative Research: Opportunities, Limitations, and Practical Considerations

DGX agent

arXiv:2605.16538v1 Announce Type: cross Abstract: This paper examines the opportunities, limitations, and practical considerations associated with the use of large language models (LLMs) in qualitativ

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

DGX agent

arXiv:2601.06943v2 Announce Type: replace-cross Abstract: In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed ac

model-releasesarxiv-cs-ai
19 May 2026
Agents

OpenAaaS: An Open Agent-as-a-Service Framework for Distributed Materials-Informatics Research

DGX agent

arXiv:2605.13618v1 Announce Type: cross Abstract: The Materials Genome Initiative catalyzed the proliferation of centralized platforms--SaaS, PaaS, and IaaS--that aggregate computational and experimen

agentsarxiv-cs-ai
14 May 2026
Local Ai

PROMETHEUS: Automating Deep Causal Research Integrating Text, Data and Models

DGX agent

arXiv:2605.12835v1 Announce Type: new Abstract: Large language models can extract local causal claims from text, but those claims become more useful when organized as persistent, navigable world model

local-aiarxiv-cs-ai
14 May 2026
Hardware

AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration -- Learning from Cheap, Optimizing Expensive

DGX agent

arXiv:2605.11518v1 Announce Type: cross Abstract: Effectively configuring scalable large language model (LLM) experiments, spanning architecture design, hyperparameter tuning, and beyond, is crucial f

hardwarearxiv-cs-cl
13 May 2026
Model Releases

Re^2Math: Benchmarking Theorem Retrieval in Research-Level Mathematics

DGX agent

arXiv:2605.09012v1 Announce Type: new Abstract: Large language models are increasingly capable at closed-world mathematical reasoning, but research assistance also requires source-grounded use of the

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

DGX agent

arXiv:2605.09063v1 Announce Type: new Abstract: Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challengi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Beating the Style Detector: Three Hours of Agentic Research on the AI-Text Arms Race

DGX agent

arXiv:2605.02620v1 Announce Type: new Abstract: Reproducing an empirical NLP study used to take weeks. Given the released data and a modern agentic-research harness, we redo every experiment of a rece

model-releasesarxiv-cs-cl
5 May 2026
Tutorials

Measuring AI Reasoning: A Guide for Researchers

DGX agent

arXiv:2605.02442v1 Announce Type: cross Abstract: In this paper, we offer a guide for researchers on evaluating reasoning in language models, building the case that reasoning should be assessed throug

tutorialsarxiv-cs-cl
5 May 2026
Applications

Pedagogical Promise and Peril of AI: A Text Mining Analysis of ChatGPT Research Discussions in Programming Education

DGX agent

arXiv:2605.00361v1 Announce Type: cross Abstract: GenAI systems such as ChatGPT are increasingly discussed in programming education, but the ways in which the research literature conceptualizes and fr

applicationsarxiv-cs-ai
5 May 2026
Model Releases

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning

DGX agent

arXiv:2605.01489v1 Announce Type: cross Abstract: Frontier scientific reasoning is rapidly emerging as a key foundation for advancing AI agents in automated scientific discovery. Deep research agents

model-releasesarxiv-cs-cl
5 May 2026
Safety

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification

DGX agent

arXiv:2601.15808v2 Announce Type: replace Abstract: Recent advances in Deep Research Agents (DRAs) are transforming automated knowledge discovery and problem-solving. While the majority of existing ef

safetyarxiv-cs-ai
30 Apr 2026
Model Releases

Beyond coauthorship: semantic structure and phantom collaborators in transportation research, 1967--2025

DGX agent

arXiv:2604.23699v1 Announce Type: cross Abstract: We present a semantic-structural atlas of transportation research built from 120{,}323 papers across 34 peer-reviewed journals published between 1967

model-releasesarxiv-cs-lg
28 Apr 2026
Model Releases

Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination

DGX agent

arXiv:2604.24690v1 Announce Type: new Abstract: While Large Language Models (LLMs) have increasingly assisted in historical tasks such as text processing, their capacity for professional-level histori

model-releasesarxiv-cs-cl
28 Apr 2026
Safety

OpenPodcar2: a robust, ROS2 vehicle for self-driving research

DGX agent

arXiv:2604.24242v1 Announce Type: new Abstract: OpenPodcar2 is a robust, ROS2-interfaced, low-cost, open source hardware and software, autonomous vehicle platform based on an off-the-shelf, hard-canop

safetyarxiv-cs-ro
28 Apr 2026
Applications

Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions

DGX agent

arXiv:2604.15395v1 Announce Type: new Abstract: Over the recent years, the field of robotics has been undergoing a transformative paradigm shift from fixed, single-task, domain-specific solutions towa

applicationsarxiv-cs-ro
20 Apr 2026
Model Releases

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

DGX agent

arXiv:2604.12102v1 Announce Type: new Abstract: We introduce compute-grounded reasoning (CGR), a design paradigm for spatial-aware research agents in which every answerable sub-problem is resolved by

model-releasesarxiv-cs-ai
15 Apr 2026
Applications

ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery

DGX agent

arXiv:2604.09237v1 Announce Type: new Abstract: Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, tradition

applicationsarxiv-cs-cl
13 Apr 2026
← Previous
1…56789…458
Next →