ToolClaude Code8 recent entries4 Aug 2026Prompt-Induced Waste in Large Reasoning Models: A Preregistered Two-Harness Benchmark of Coding AgentsarXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wordin→4 Aug 2026LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding AgentarXiv:2608.00267v1 Announce Type: cross Abstract: Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon soft
ToolCursor3 recent entries12 May 2026Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure VariablesarXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid→27 May 2026Learning GUI Grounding with Spatial Reasoning from Visual FeedbackarXiv:2509.21552v2 Announce Type: replace-cross Abstract: Graphical User Interface (GUI) grounding is commonly framed as a coordinate prediction task -- given a natural language instruction, generate →30 Jul 2026(Im)Paired Programming: Coding Agents Improve Productivity but Harm UnderstandingarXiv:2607.26375v1 Announce Type: new Abstract: Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewi
ToolLangChain3 recent entries12 Apr 2026Wiki Lint Report — 2026-04-12Automated lint: 34 errors, 0 warnings, 3 info→15 Apr 2026OctoTools: An Agentic Framework with Extensible Tools for Complex ReasoningarXiv:2502.11271v2 Announce Type: replace-cross Abstract: Solving complex reasoning tasks may involve visual understanding, domain knowledge retrieval, numerical calculation, and multi-step reasoning.→26 Apr 2026Wiki Lint Report — 2026-04-26Automated lint: 44 errors, 10 warnings, 3 info
ToolOllama6 recent entries28 Apr 2026Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter ScalearXiv:2604.23801v1 Announce Type: new Abstract: Practitioners deploying small open-weight large language models (LLMs) for medical question answering face a recurring design choice: invest in a domain→21 May 2026GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema RetrievalarXiv:2605.20815v1 Announce Type: new Abstract: Graph-based Retrieval Augmented Generation (GraphRAG) extends retrieval-augmented generation to support structured reasoning over complex corpora, but i→1 Jun 2026Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation WorkflowsarXiv:2605.31452v1 Announce Type: new Abstract: Building on our previous work, this paper develops practical, low-barrier methods for freelance translators and smaller language service providers to ev→2 Jun 2026Benchmarking Local LLMs for Natural-Language-to-SQL Querying in Biopharmaceutical Manufacturing: An Empirical Benchmark on Consumer-Grade HardwarearXiv:2606.01338v1 Announce Type: new Abstract: Biopharmaceutical manufacturing organizations operate under regulatory frameworks such as FDA guidance, EU Good Manufacturing Practice (GMP), and the EU→10 Jul 2026SPL: Orchestrating Workflows with Declarative Deterministic-Probabilistic CompositionarXiv:2607.07727v1 Announce Type: cross Abstract: We present SPL (Structured Prompt Language), a declarative language that composes deterministic and probabilistic computation modes in a single specif→28 Jul 2026Retrieval-Augmented Large Language Models as Components of Cognitive Computing architecture for Regulatory Knowledge ManagementarXiv:2607.24352v1 Announce Type: new Abstract: The aim of this article is to verify whether integrating large language models (LLMs) with the Retrieval-Augmented Generation (RAG) architecture enables
ToolVercel AI2 recent entries17 Apr 2026Dive into Claude Code: The Design Space of Today's and Future AI Agent SystemsarXiv:2604.14228v1 Announce Type: cross Abstract: Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study describes →7 Jul 2026When Words Predict WorkloadarXiv:2607.04951v1 Announce Type: cross Abstract: Standard distributed ac{llm} schedulers rely on static token counts or rolling latency averages, making them susceptible to failures on statutorily co
ToolHugging Face8 recent entries24 Jun 2026SciZoom: A Large-scale Benchmark for Hierarchical Scientific Summarization across the LLM EraarXiv:2603.16131v2 Announce Type: replace Abstract: The explosive growth of AI research has created unprecedented information overload, increasing the demand for scientific summarization at multiple l→30 Jun 2026Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African LanguagesarXiv:2606.28867v1 Announce Type: new Abstract: Creative Commons licenses dominate African NLP corpus releases, but their compatibility rules are rarely applied. CC-BY-SA and CC-BY-NC cannot be combin→1 Jul 2026BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance ExamsarXiv:2606.22723v2 Announce Type: replace Abstract: Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended, →24 Jul 2026Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization ModelsarXiv:2607.20442v1 Announce Type: new Abstract: We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in Jul→24 Jul 2026GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction CorpusarXiv:2607.20443v1 Announce Type: new Abstract: We release GLAN-QnA-KR, a 303,581-row openly redistributable Korean instruction-QA corpus produced via the seedless taxonomy-driven GLAN synthesis pipel→24 Jul 2026DONDO: Open w2v-BERT Speech-Recognition Base Models for African LanguagesarXiv:2607.21540v1 Announce Type: new Abstract: We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.→11 Aug 2026Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and RerankersarXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the rig→12 Aug 2026myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASRarXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work prese