AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “engineering”

GridTimelineEvolution
3,109 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
4 Aug 2026LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent

arXiv:2608.00267v1 Announce Type: cross Abstract: Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon soft

→5 Aug 2026VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space

arXiv:2608.02878v1 Announce Type: new Abstract: Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on stan


Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→5 Aug 2026Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation

arXiv:2608.02672v1 Announce Type: cross Abstract: Cloud misconfiguration remains a leading cause of security incidents, yet whether LLMs and SLMs can generate security-compliant Infrastructure-as-Code

→6 Aug 2026Transfer Learning for Named Entity Recognition of Classical Latin through LLM Prompting

arXiv:2608.04015v1 Announce Type: new Abstract: With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient lang

→6 Aug 2026One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIP

arXiv:2505.19840v3 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) have achieved widespread success yet remain prone to adversarial attacks. Typically, such attacks either involve f

→6 Aug 2026Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools

arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Mo

→11 Aug 2026LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs

arXiv:2608.08467v1 Announce Type: new Abstract: The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequentl

→12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses

arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these

CompanyOpenAI8 recent entries
10 Jul 2026Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

arXiv:2607.08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, t

→15 Jul 2026Optimization Is Not All You Need

arXiv:2607.11977v1 Announce Type: new Abstract: In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produce

→15 Jul 2026Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain

arXiv:2607.07652v1 Announce Type: cross Abstract: Search engines have long allocated attention on the web by routing users from queries to websites. AI search changes this arrangement because informat

→24 Jul 2026InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

arXiv:2607.20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or n

→29 Jul 2026Many-body Tipping Dynamics of ChatGPT-like AIs

arXiv:2607.25279v1 Announce Type: new Abstract: Why do ChatGPT-like AIs, despite major architectural and training differences, unexpectedly tip to undesirable content (e.g. harmful, misleading, repeti

→10 Aug 2026Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks

arXiv:2608.07283v1 Announce Type: new Abstract: Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cr

→11 Aug 2026Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning

arXiv:2608.07593v1 Announce Type: cross Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Exi

→11 Aug 2026Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks

arXiv:2512.22966v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision

CompanyGoogle8 recent entries
24 Jul 2026AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs

arXiv:2607.20498v1 Announce Type: new Abstract: Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-h

→5 Aug 2026Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation

arXiv:2608.02672v1 Announce Type: cross Abstract: Cloud misconfiguration remains a leading cause of security incidents, yet whether LLMs and SLMs can generate security-compliant Infrastructure-as-Code

→6 Aug 2026Transfer Learning for Named Entity Recognition of Classical Latin through LLM Prompting

arXiv:2608.04015v1 Announce Type: new Abstract: With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient lang

→6 Aug 2026The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning

arXiv:2608.04285v1 Announce Type: new Abstract: Neurosymbolic AI systems that integrate machine learning and symbolic reasoning are rapidly gaining attention. They complement the data-intensive statis

→6 Aug 2026NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts

arXiv:2608.04030v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domains remains

→7 Aug 2026Vibe Compiler: A Research-Logic Synthesis Tool That Runs without Prompt Engineering -Toward Enhancing Metacognition for Sustaining Agency in the Age of Generative AI-

arXiv:2608.05545v1 Announce Type: cross Abstract: Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by encouraging unc

→11 Aug 2026Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks

arXiv:2512.22966v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision

→11 Aug 2026LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs

arXiv:2608.08467v1 Announce Type: new Abstract: The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequentl

CompanyMeta8 recent entries
6 Aug 2026Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools

arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Mo

→10 Aug 2026Latent Fact-Checking: Detecting Misinformation through Activation Engineering

arXiv:2608.06417v1 Announce Type: cross Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level ling

→10 Aug 2026A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy

arXiv:2608.07427v1 Announce Type: new Abstract: LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network anal

→11 Aug 2026LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

arXiv:2608.08382v1 Announce Type: new Abstract: As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control. Enabling fractio

→11 Aug 2026Large Language Models Align with the Human Brain during Creative Thinking

arXiv:2604.03480v2 Announce Type: replace-cross Abstract: Creative thinking is a fundamental aspect of human cognition, and divergent thinking-the capacity to generate novel and varied ideas-is widely

→11 Aug 2026Contamination Means Overestimation? A Fine-Grained Empirical Study in Code Intelligence

arXiv:2506.02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread

→12 Aug 2026SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents

arXiv:2608.10157v1 Announce Type: new Abstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over t

→12 Aug 2026Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic Systems

arXiv:2608.10941v1 Announce Type: new Abstract: Industrial time-series signals, such as turbine temperature and rotational speed in aero-engines, are essential for monitoring the health and operationa

CompanyMistral8 recent entries
11 May 2026CommandSwarm: Safety-Aware Natural Language-to-Behavior-Tree Generation for Robotic Swarms

arXiv:2605.07764v1 Announce Type: new Abstract: Natural-language interfaces can make swarm robotics more accessible to non-expert operators, but they must translate ambiguous user intent into executab

→4 Jun 2026Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs

arXiv:2606.04035v1 Announce Type: cross Abstract: We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5

→8 Jun 2026Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs

arXiv:2601.12359v1 Announce Type: cross Abstract: Prompt injection attacks have become an increasing vulnerability for LLM applications, where adversarial prompts exploit indirect input channels such

→30 Jun 2026Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study

arXiv:2606.29213v1 Announce Type: new Abstract: OCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on

→28 Jul 2026The Few-shot Dilemma: Over-prompting Large Language Models

arXiv:2509.13196v2 Announce Type: replace Abstract: Over-prompting, a phenomenon where excessive examples in prompts lead to diminished performance in Large Language Models (LLMs), challenges the conv

→28 Jul 2026LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph

arXiv:2607.24551v1 Announce Type: new Abstract: Maintenance regulations are complex legal texts that are difficult to exploit when addressing a specific case and challenging to integrate into operatio

→29 Jul 2026Influence of Prompt Engineering on Small Language Models for Guarded Query Routing

arXiv:2607.24801v1 Announce Type: cross Abstract: We study the problem of guarded query routing, where we assume that a user query first meets a router that either determines the ideal endpoint for in

→5 Aug 2026Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform

arXiv:2608.03311v1 Announce Type: cross Abstract: Enterprise compliance management requires rapid adaptation to evolving regulatory frameworks (e.g., DORA, AI RMF, FedRAMP) and tight remediation SLAs.

CompanyxAI8 recent entries
28 Apr 2026Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems

arXiv:2602.04120v2 Announce Type: replace-cross Abstract: Though Explainable AI (XAI) has made significant advancements, its inclusion in edge and IoT systems is typically ad-hoc and inefficient. Most

→28 Apr 2026Explainable Artificial Intelligence Techniques for Interpretation of Food Models: a Review

arXiv:2504.10527v2 Announce Type: replace Abstract: Artificial Intelligence (AI) has become essential for analyzing complex data and solving highly-challenging tasks. It is being applied across numero

→5 Jun 2026Grounded but Misleading: Evaluating Semantic Alignment in AI-Generated Security Explanations

arXiv:2602.05056v2 Announce Type: replace-cross Abstract: Online scams increasingly leverage fluent and context-aware social engineering strategies, creating growing demand for AI systems that explain

→9 Jun 2026Impacts of Histories and Models on LLM Grading: A Study in Advanced Software Engineering Courses

arXiv:2606.08400v1 Announce Type: cross Abstract: Graduate-level research reading report assessment creates a substantial labor burden for educators. While large language models (LLMs) hold great pote

→1 Jul 2026Partition-Guided Distance Saliency: Bridging Decision and Objective Spaces in Many-Objective Optimization

arXiv:2606.30836v1 Announce Type: new Abstract: Explainability in Many-Objective Optimization (MaO) is currently hindered by the escalating complexity of the Pareto front, which renders the relationsh

→9 Jul 2026Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic

→15 Jul 2026Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability

arXiv:2607.12056v1 Announce Type: new Abstract: Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and

→31 Jul 2026AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design Stages

arXiv:2505.10300v2 Announce Type: replace-cross Abstract: Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle through

CompanyDeepSeek8 recent entries
7 Jul 2026Dictionaries, Not Darwin: Set-Level Selection Beats LLM Evolution in Scientific Equation Discovery

arXiv:2607.04108v1 Announce Type: new Abstract: Large language models are increasingly used as evolutionary engines for scientific discovery: generate candidates, select winners, feed them back as par

→8 Jul 2026Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models

arXiv:2607.06012v1 Announce Type: new Abstract: Real estate property listings expose structured metadata through the API. Still, the richest property-level information (i.e., legal status, structural

→8 Jul 2026Foundation Models for Automatic CAD Generation

arXiv:2607.05573v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) enable the automatic generation of parametric 3D designs from natural-

→10 Jul 2026Infinity-Parser2 Technical Report

arXiv:2607.07836v1 Announce Type: new Abstract: We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end

→10 Jul 2026Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment

arXiv:2607.08268v1 Announce Type: new Abstract: High-volume structured extraction pays a large model's latency on every item, so distilling the task into a small on-device model is attractive: compara

→28 Jul 2026The Few-shot Dilemma: Over-prompting Large Language Models

arXiv:2509.13196v2 Announce Type: replace Abstract: Over-prompting, a phenomenon where excessive examples in prompts lead to diminished performance in Large Language Models (LLMs), challenges the conv

→31 Jul 2026Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

arXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an

→31 Jul 2026A Physics-Informed Framework for PID Tuning of Chemical Processes Using Large Language Model Agents

arXiv:2607.26594v1 Announce Type: cross Abstract: PID tuning for chemical processes commonly relies on identified process models, whereas plant engineers often retune loops iteratively by observing re

CompanyNVIDIA8 recent entries
31 Jul 2026KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation

arXiv:2607.27231v1 Announce Type: cross Abstract: Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specia

→4 Aug 2026Open-DiffLoco: Open-Source Differentiable Learning for Deployable Blind Quadruped Locomotion

arXiv:2608.02069v1 Announce Type: cross Abstract: Developing deployable locomotion policies through conventional reinforcement learning often requires complex reward engineering and expensive training

→6 Aug 2026RORA: Realistic Object Reconstruction with Articulation

arXiv:2608.04842v1 Announce Type: new Abstract: Replicating real-world environments into simulation by realistic visual representation like NeRF and 3D Gaussian Splatting (3DGS) has emerged as an effe

→6 Aug 2026An Explainable LLM Agent Layer for Open-World Anomaly Detection in Oil Wells

arXiv:2608.04041v1 Announce Type: new Abstract: Open-World Learning (OWL) pipelines for oil well anomaly detection have recently been shown to combine autoencoder-based detection, multiclass classific

→7 Aug 2026IcFuzz: Fuzzing Isaac Sim with Semantic Stage Guidance and Multi-level Mutation

arXiv:2608.06088v1 Announce Type: new Abstract: Robotics simulators serve as a foundational infrastructure for embodied AI, facilitating safe and scalable robotic system development. NVIDIA Isaac Sim

→10 Aug 2026Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control

arXiv:2608.06758v1 Announce Type: new Abstract: We present Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built from Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16

→10 Aug 2026Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-Tuning

arXiv:2608.07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class

→12 Aug 2026HyWA: Architecture-Preserving Personalized Voice Activity Detection for Full-Duplex Voice Assistants

arXiv:2510.12947v3 Announce Type: replace-cross Abstract: Voice activity detection (VAD) serves as an early gate in voice-assistant pipelines for smart devices. Because conventional VADs respond to sp