AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “engineering”

GridTimelineEvolution
5,413 results
6 Aug 2026

Auto-fit vs tuned MoE offload: 564 → 1330 pp tok/s, unchanged decode (Qwen3.6-35B-A3B Q6 / RTX 3090)

Model ReleasesDGX agent

TL;DR: On a Qwen3.6-35B-A3B Q6 setup sized for 64K context on a 24GB RTX 3090, spilling eight MoE expert layers to CPU freed enough VRAM to increase -b from 512 to 1024 and -ub from 128 to 512. Prompt

Best open-source harnesses for combining cloud and local AI model orchestration?

Local AiDGX agent

Looking for best current solutions for combining cloud models and local models seamlessly inside a harness' orchestration Edit: Right now, we don't have harnesses (that I'm aware of) that are blending

CoCo-IR: Contextual Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2608.05149v1 Announce Type: new Abstract: Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of compl


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Continual-Learning Physics-Informed Neural Networks for Parameterized Partial Differential Equations

Model ReleasesDGX agent

arXiv:2608.04778v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) incorporate governing equations into neural-network training and can approximate PDE solutions without requirin

DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation

SafetyDGX agent

arXiv:2608.04622v1 Announce Type: new Abstract: AI agents have emerged as a powerful new paradigm in generative image synthesis, enabling systems to perform complex semantic reasoning rather than pass

DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

Model ReleasesDGX agent

DeepSeek‑V4 Flash 0731 is the cheapest model on the DeepSWE board, costing about 0.10 per rollout versus GPT‑5.6 Luna’s 0.61, yet it scores a pass@1 of 53.3% compared to Luna’s 67.2%. A cascade strate

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools

Model ReleasesDGX agent

arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Mo

EdgeLM: Edge Demonstrations for Language Models' Table Understanding

ResearchDGX agent

arXiv:2608.04390v1 Announce Type: new Abstract: Large language models (LLMs) perform table-centric prediction through in-context learning, making demonstration selection critical to performance. Exist

Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning

AgentsDGX agent

arXiv:2608.04457v1 Announce Type: cross Abstract: As 'AI Scientists' emerge to drive research via the Model Context Protocol (MCP), systems relying on ephemeral scripts will fail. The sheer scale of s

EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement

Local AiDGX agent

arXiv:2608.04968v1 Announce Type: new Abstract: The capabilities of an LLM agent depend not only on its model but on the harness: the executable program that constructs context, invokes tools, verifie

FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

Model ReleasesDGX agent

arXiv:2608.04077v1 Announce Type: new Abstract: Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompt

From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs

SafetyDGX agent

arXiv:2601.03808v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved notable performance in code synthesis; however, data-aware augmentation remains a limiting factor, handle

From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking

Model ReleasesDGX agent

arXiv:2608.05030v1 Announce Type: new Abstract: Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, ex

I don't understand what's controversial here. Books should be written as they have always been. With ink made from crushing the bodies of ra…

AgentsDGX agent

I don't understand what's controversial here. Books should be written as they have always been. With ink made from crushing the bodies of rare insects from Turkey you harvested yourself and then groun

i just spent weeks rewriting my webUI from scratch, getting rid of all AI slop within the codebase and switching it over to a proper lightweight framework (alpine.js). i am now comfortable suggesting it as an alternative to openwebUI, librechat and the like! it is made for local models

Local AiDGX agent

[Fully open source under GPL3, made from the ground up for use with local models, no subscriptions, no corporate backing] When i first started this, it was meant to be a fully lightweight, extremely m

Learning Compression Rules for Network Traffic

ApplicationsDGX agent

arXiv:2608.04545v1 Announce Type: new Abstract: We study the problem of learning compact rule-based compressors for structured network traffic. Each packet is a record of header fields that are highly

MIDAS: Multi-LLM Iterative Data-Adaptive Summarization

ApplicationsDGX agent

arXiv:2608.04307v1 Announce Type: cross Abstract: Text summarization is deceptively difficult. While condensing information seems straightforward, real-world enterprise summarization of support ticket

Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications

Model ReleasesDGX agent

Nearly every major AI lab uses Google Cloud infrastructure, including for training of models, inference for agents, and new frontier research. Google Cloud also continues to be the platform of choice

OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

AgentsDGX agent

arXiv:2608.05141v1 Announce Type: new Abstract: Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon

OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing

Model ReleasesDGX agent

arXiv:2608.04434v1 Announce Type: new Abstract: Recent large language models (LLMs) have demonstrated remarkable progress in constraint-aware navigation, maze reasoning, and graph reasoning. However,

One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIP

Model ReleasesDGX agent

arXiv:2505.19840v3 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) have achieved widespread success yet remain prone to adversarial attacks. Typically, such attacks either involve f

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories

AgentsDGX agent

arXiv:2604.07341v2 Announce Type: replace-cross Abstract: Most repository-level code translation and validation techniques have been evaluated on a single source-target programming language (PL) pair,

RORA: Realistic Object Reconstruction with Articulation

HardwareDGX agent

arXiv:2608.04842v1 Announce Type: new Abstract: Replicating real-world environments into simulation by realistic visual representation like NeRF and 3D Gaussian Splatting (3DGS) has emerged as an effe

SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts

SafetyDGX agent

arXiv:2608.04962v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves the reasoning capabilities of large language models, but autoregressive rollout generation remains

The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning

TutorialsDGX agent

arXiv:2608.04285v1 Announce Type: new Abstract: Neurosymbolic AI systems that integrate machine learning and symbolic reasoning are rapidly gaining attention. They complement the data-intensive statis

Transfer Learning for Named Entity Recognition of Classical Latin through LLM Prompting

Model ReleasesDGX agent

arXiv:2608.04015v1 Announce Type: new Abstract: With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient lang

TriCLE: Tri-Modal Vision-Language Reasoning for Edge-Deployed Fine-Grained Clustering

SafetyDGX agent

arXiv:2608.04175v1 Announce Type: new Abstract: Edge platforms used for aerial observation must interpret aircraft imagery under limited memory, limited compute, and intermittent connectivity. This se

TS2TabPFN: Time Series Classification and Extrinsic Regression through Feature Extraction and a Tabular Foundation Model

ResearchDGX agent

arXiv:2608.04174v1 Announce Type: new Abstract: Time series data are ubiquitous in practical applications, where classification (TSC) and extrinsic regression (TSER) have emerged as essential tasks fo

War in the Abstract: The Rise and Consequences of Militarized Language in Scientific Communication

SafetyDGX agent

arXiv:2606.23462v2 Announce Type: replace-cross Abstract: Scientists do not, by profession, wage war. Yet warfare's vocabulary consistently appears in their abstracts. To quantify the extent to which

Your agentic summer: No-cost lessons from Google experts to build and scale agents

Model ReleasesDGX agent

I’ve talked to developers, IT leaders, and builders who all ask the same question: How do we actually get agents into production? The answer isn't theoretical — it's hands-on. Whether it’s designing a

Zero-shot reasoning for simulating scholarly peer-review

Model ReleasesDGX agent

arXiv:2510.02027v2 Announce Type: replace Abstract: Scholarly publishing requires scalable scrutiny supported by auditable evidence. This paper presents a two-component benchmark of xPeer, the peer-re

5 Aug 2026

A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

SafetyDGX agent

arXiv:2608.02684v1 Announce Type: cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in

AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits

Model ReleasesDGX agent

arXiv:2608.03738v1 Announce Type: new Abstract: As Moore's law slows, the industry is turning to three-dimensional integration; yet in merged 3D-IC flows, routed designs expose bond-level defects with

AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection

Local AiDGX agent

arXiv:2601.19138v2 Announce Type: replace-cross Abstract: Secure code review is critical during pre-integration, where Atlassian developers rely on lightweight analysis tools, while deep security asse

Announcing Discovery Loop! I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolViny…

TutorialsDGX agent

Announcing Discovery Loop! I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoL

Calibrating Semantic Uncertainty from Observable Language-Model Probabilities

ResearchDGX agent

arXiv:2607.17447v2 Announce Type: replace-cross Abstract: As generative artificial intelligence enters scientific and professional work, its uncertainty must be defined on the states that matter for i

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning

SafetyDGX agent

arXiv:2608.03875v1 Announce Type: cross Abstract: Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Mode

Equivariant Music Transformer

ResearchDGX agent

arXiv:2608.03920v1 Announce Type: cross Abstract: Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation s

Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform

Model ReleasesDGX agent

arXiv:2608.03311v1 Announce Type: cross Abstract: Enterprise compliance management requires rapid adaptation to evolving regulatory frameworks (e.g., DORA, AI RMF, FedRAMP) and tight remediation SLAs.

Harness choice is a big deal. So much room to advance and improve results across the board with agent harnesses. Great paper highlighting th…

Model ReleasesDGX agent

Harness choice is a big deal. So much room to advance and improve results across the board with agent harnesses. Great paper highlighting this. New research releases DataSpace, a benchmark where data

Incident Report: unsanctioned agent behaviour during cyber testing

Model ReleasesDGX agent

Incident Report: unsanctioned agent behaviour during cyber testing It happened again. This time it was the UK government's AI Security Institute who accidentally attacked other companies while running

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

AgentsDGX agent

arXiv:2608.03644v1 Announce Type: new Abstract: AI agents deployed in real-world settings must be capable of coordinating with humans and other AI agents they have not encountered before. Zero-shot co

LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds

Model ReleasesDGX agent

arXiv:2608.03078v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong defect recognition capability in industrial anomaly detection. However, in lithography review,

MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

Model ReleasesDGX agent

TensorSharp's MoE CPU-offload feature has been merged into main. Here is the parameters description of this feature: Mixture-of-Experts CPU offload: --n-cpu-moe <N> | -ncmoe <N> Keep the routed MoE ex

MutMem: Cryptographically Authorized Mutation in Persistent Agent Memory

SafetyDGX agent

arXiv:2608.02843v1 Announce Type: cross Abstract: Persistent agent memory must adapt as later outcomes change earlier evidence, yet mutable retrieval weights create an attribution problem: reviewers m

One-shotting a Raccoon Heist game using Claude Fable 5

Model ReleasesDGX agent

Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept 'art' created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5

Open-Linguistic Concept Unified Learning for Cross-Site Interpretable Dermatology Image Diagnosis

ResearchDGX agent

arXiv:2608.03225v1 Announce Type: new Abstract: Human-interpretable computer-aided diagnosis is crucial for clinical decision making. Concept-based models excel by providing transparent reasoning and

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

SafetyDGX agent

arXiv:2608.03682v1 Announce Type: new Abstract: Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, a

Principles of Robot Autonomy

AgentsDGX agent

arXiv:2608.03496v1 Announce Type: cross Abstract: Autonomous robots are moving rapidly from research labs into everyday life - on roads, in the air, in warehouses, and in space. Robot autonomy is no l

Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation

Model ReleasesDGX agent

arXiv:2608.02672v1 Announce Type: cross Abstract: Cloud misconfiguration remains a leading cause of security incidents, yet whether LLMs and SLMs can generate security-compliant Infrastructure-as-Code

Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations

ResearchDGX agent

arXiv:2608.03970v1 Announce Type: new Abstract: Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, di

Skill libraries are shipping in agent harnesses on the assumption that writing skills down compounds. A new benchmark tests that directly. C…

Model ReleasesDGX agent

Skill libraries are shipping in agent harnesses on the assumption that writing skills down compounds. A new benchmark tests that directly. ContinualSkillBench covers five domains, each with 100 interc

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI…

SafetyDGX agent

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI model framework. I'm not surprised at all, and it's actuall

Stuck on 'A': Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model

Model ReleasesDGX agent

arXiv:2608.02689v1 Announce Type: new Abstract: We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU budg

The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems

SafetyDGX agent

arXiv:2608.03214v1 Announce Type: new Abstract: Large language models have transformed artificial intelligence from isolated prediction services into components of long-running, distributed systems th

VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space

Model ReleasesDGX agent

arXiv:2608.02878v1 Announce Type: new Abstract: Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on stan

4 Aug 2026

A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense

AgentsDGX agent

arXiv:2608.00583v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is meant to catch the reward hacks that look clean in the actions and betray themselves only in the reasoning. We sh

Beyond Random Partitioning: Unsupervised Spatio-Temporal Stratification for Cohort Balancing in Longitudinal Medical Imaging

SafetyDGX agent

arXiv:2608.00073v1 Announce Type: new Abstract: Rigorous dataset partitioning is a foundational, yet frequently overlooked, prerequisite for reliable deep learning in longitudinal medical imaging. Nai

bFaaaP: An Inclusive, Head-Angle Piano-Pedal Interaction that Quantitatively Reproduces a Pianist's Intended Pedalling -- Foot-Free, for Acoustic and Electronic Pianos

Local AiDGX agent

arXiv:2608.00633v1 Announce Type: cross Abstract: Expressive piano performance depends on the sustain (damper) pedal, operated by foot, excluding players who cannot readily use their feet: wheelchair

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

HardwareDGX agent

arXiv:2608.01651v1 Announce Type: cross Abstract: Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autore

← Previous
1…4748495051…91
Next →