AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
Human
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Tools

ToolClaude Code8 recent entries
4 Aug 2026Prompt-Induced Waste in Large Reasoning Models: A Preregistered Two-Harness Benchmark of Coding Agents

arXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wordin

→4 Aug 2026LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent

arXiv:2608.00267v1 Announce Type: cross Abstract: Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon soft

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→
4 Aug 2026DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text

arXiv:2608.01046v1 Announce Type: new Abstract: The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation

→7 Aug 2026Who Checks the Citations? Benchmarking Legal Hallucination Detection

arXiv:2606.21155v2 Announce Type: replace Abstract: Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictio

→7 Aug 2026M^3R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

arXiv:2608.05817v1 Announce Type: new Abstract: Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visu

→11 Aug 2026DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?

arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code

→12 Aug 2026VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world

→12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses

arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these

ToolCursor3 recent entries
12 May 2026Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables

arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid

→27 May 2026Learning GUI Grounding with Spatial Reasoning from Visual Feedback

arXiv:2509.21552v2 Announce Type: replace-cross Abstract: Graphical User Interface (GUI) grounding is commonly framed as a coordinate prediction task -- given a natural language instruction, generate

→30 Jul 2026(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding

arXiv:2607.26375v1 Announce Type: new Abstract: Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewi

ToolLangChain3 recent entries
12 Apr 2026Wiki Lint Report — 2026-04-12

Automated lint: 34 errors, 0 warnings, 3 info

→15 Apr 2026OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

arXiv:2502.11271v2 Announce Type: replace-cross Abstract: Solving complex reasoning tasks may involve visual understanding, domain knowledge retrieval, numerical calculation, and multi-step reasoning.

→26 Apr 2026Wiki Lint Report — 2026-04-26

Automated lint: 44 errors, 10 warnings, 3 info

ToolOllama6 recent entries
28 Apr 2026Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter Scale

arXiv:2604.23801v1 Announce Type: new Abstract: Practitioners deploying small open-weight large language models (LLMs) for medical question answering face a recurring design choice: invest in a domain

→21 May 2026GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval

arXiv:2605.20815v1 Announce Type: new Abstract: Graph-based Retrieval Augmented Generation (GraphRAG) extends retrieval-augmented generation to support structured reasoning over complex corpora, but i

→1 Jun 2026Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows

arXiv:2605.31452v1 Announce Type: new Abstract: Building on our previous work, this paper develops practical, low-barrier methods for freelance translators and smaller language service providers to ev

→2 Jun 2026Benchmarking Local LLMs for Natural-Language-to-SQL Querying in Biopharmaceutical Manufacturing: An Empirical Benchmark on Consumer-Grade Hardware

arXiv:2606.01338v1 Announce Type: new Abstract: Biopharmaceutical manufacturing organizations operate under regulatory frameworks such as FDA guidance, EU Good Manufacturing Practice (GMP), and the EU

→10 Jul 2026SPL: Orchestrating Workflows with Declarative Deterministic-Probabilistic Composition

arXiv:2607.07727v1 Announce Type: cross Abstract: We present SPL (Structured Prompt Language), a declarative language that composes deterministic and probabilistic computation modes in a single specif

→28 Jul 2026Retrieval-Augmented Large Language Models as Components of Cognitive Computing architecture for Regulatory Knowledge Management

arXiv:2607.24352v1 Announce Type: new Abstract: The aim of this article is to verify whether integrating large language models (LLMs) with the Retrieval-Augmented Generation (RAG) architecture enables

ToolVercel AI2 recent entries
17 Apr 2026Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems

arXiv:2604.14228v1 Announce Type: cross Abstract: Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study describes

→7 Jul 2026When Words Predict Workload

arXiv:2607.04951v1 Announce Type: cross Abstract: Standard distributed ac{llm} schedulers rely on static token counts or rolling latency averages, making them susceptible to failures on statutorily co

ToolHugging Face8 recent entries
24 Jun 2026SciZoom: A Large-scale Benchmark for Hierarchical Scientific Summarization across the LLM Era

arXiv:2603.16131v2 Announce Type: replace Abstract: The explosive growth of AI research has created unprecedented information overload, increasing the demand for scientific summarization at multiple l

→30 Jun 2026Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages

arXiv:2606.28867v1 Announce Type: new Abstract: Creative Commons licenses dominate African NLP corpus releases, but their compatibility rules are rarely applied. CC-BY-SA and CC-BY-NC cannot be combin

→1 Jul 2026BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

arXiv:2606.22723v2 Announce Type: replace Abstract: Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended,

→24 Jul 2026Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization Models

arXiv:2607.20442v1 Announce Type: new Abstract: We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in Jul

→24 Jul 2026GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus

arXiv:2607.20443v1 Announce Type: new Abstract: We release GLAN-QnA-KR, a 303,581-row openly redistributable Korean instruction-QA corpus produced via the seedless taxonomy-driven GLAN synthesis pipel

→24 Jul 2026DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

arXiv:2607.21540v1 Announce Type: new Abstract: We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.

→11 Aug 2026Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

arXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the rig

→12 Aug 2026myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

arXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work prese