AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,762 results
3 Aug 2026

two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apo…

Model ReleasesDGX agent

two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apology. i'm sorry that i was right about every single thing. a

2 Aug 2026

All other models on the Portal remain 20% discounted, aside from GPT-5.6 Terra and Luna which are 50% off. https://x.com/NousResearch/status…

Model ReleasesDGX agent

All other models on the Portal remain 20% discounted, aside from GPT-5.6 Terra and Luna which are 50% off. https://x.com/NousResearch/status/2080039066771337475?s=20 All models are now 20% off for a l

Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization - AI's narrative

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

# Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization I used Deepseek-v4-Flash-0731 cloud API settig up vllm-moet to run deepseek-v4-flash with MTP locally on

To address the limits of deep learning and avoid stalling, the field of AI started by applying patch (1), which started being demoed 9 month…

ResearchDGX agent

To address the limits of deep learning and avoid stalling, the field of AI started by applying patch (1), which started being demoed 9 months later in December 2024 and has now become completely ubiqu

31 Jul 2026

A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response

SafetyDGX agent

arXiv:2607.27597v1 Announce Type: new Abstract: Recent advances in Vision Language Models (VLMs) have created new opportunities for disaster response, where responders must interpret large volumes of

AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching

Model ReleasesDGX agent

arXiv:2607.27130v1 Announce Type: new Abstract: Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

SafetyDGX agent

arXiv:2607.26639v1 Announce Type: cross Abstract: A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% def

Bridging AI and Energy Forecasting: An Autonomous Workflow with Customized Toolkit

Model ReleasesDGX agent

arXiv:2307.07191v3 Announce Type: replace Abstract: Energy forecasting is crucial for the power grid, but fundamentally different from general time series analysis: it highly relies on covariates like

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

Model ReleasesDGX agent

arXiv:2607.27747v1 Announce Type: new Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either

Can Vision-Language Models Reason about AI Edits in Images?

Local AiDGX agent

arXiv:2607.28464v1 Announce Type: new Abstract: Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly

Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy

SafetyDGX agent

arXiv:2607.27212v1 Announce Type: cross Abstract: Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of q

DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

Model ReleasesDGX agent

arXiv:2607.27263v1 Announce Type: new Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

Model ReleasesDGX agent

arXiv:2607.26611v1 Announce Type: new Abstract: AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in us

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Model ReleasesDGX agent

arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of do

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

Model ReleasesDGX agent

Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without

LLMs struggle to simulate human belief updates in controlled environments

Model ReleasesDGX agent

arXiv:2607.28347v1 Announce Type: new Abstract: LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

HardwareDGX agent

arXiv:2607.28312v1 Announce Type: new Abstract: Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches prima

Optimal Realistic Local AI for Most

Model ReleasesDGX agent

So you’ve got a 3090 or maybe even a 5090? Or more likely a 4060 8GB Ti. You wanna try local AI, you don’t know what it can/can’t do. 1) Install the best model you can. If you have a 3090 or a 5090, t

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

Model ReleasesDGX agent

arXiv:2607.28008v1 Announce Type: new Abstract: Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthe

RLPF: Reinforcement Learning from Performance Feedback for Code Generation

TutorialsDGX agent

arXiv:2607.27271v1 Announce Type: new Abstract: Code models are increasingly trained with execution feedback, but most training signals still stop at correctness. This leaves an important gap for syst

Scaling medical imaging report generation with multimodal reinforcement learning

Model ReleasesDGX agent

arXiv:2601.17151v2 Announce Type: replace-cross Abstract: Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit ma

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

Model ReleasesDGX agent

arXiv:2607.27421v1 Announce Type: new Abstract: Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

SafetyDGX agent

arXiv:2607.26566v1 Announce Type: cross Abstract: Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them inter

Simulation of Surgical Suturing Using Position-Based Dynamics and the Material Point Method for Robot Reinforcement Learning

HardwareDGX agent

arXiv:2607.27494v1 Announce Type: new Abstract: Recent advances in robotics research have created a strong demand for high-performance simulators. Surgical robotics simulation faces unique challenges

smevals - a small eval suite for evaluating models, prompts, and harnesses

Model ReleasesDGX agent

smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer

THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model

Model ReleasesDGX agent

arXiv:2607.27303v1 Announce Type: cross Abstract: Temporal heterogeneous graphs offer a natural abstraction for dynamic relational systems in which diverse node and relation types co-exist and evolve

Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

Model ReleasesDGX agent

arXiv:2607.28287v1 Announce Type: cross Abstract: ARC-AGI-3 turns abstraction into an interactive problem of skill acquisition. A player must infer an unfamiliar game's rules, hidden state, and goal w

Very interesting paper on recursive self-improvement. The whole stack is released. Machine learning engineering gives recursive self-improve…

Model ReleasesDGX agent

Very interesting paper on recursive self-improvement. The whole stack is released. Machine learning engineering gives recursive self-improvement a concrete, executable testbed. OpenMLE is an open full

30 Jul 2026

APEX-Accounting

Model ReleasesDGX agent

arXiv:2607.27189v1 Announce Type: new Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountant

ARC-Encoder: learning compressed text representations for large language models

Model ReleasesDGX agent

arXiv:2510.20535v2 Announce Type: replace Abstract: Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs. Co

Batten Down Your Packages: Mitigation Guidance for Supply Chain Compromise

Model ReleasesDGX agent

Written by: Kelli Vanderlee, Stuart Carrera For years, the cybersecurity industry's understanding of software supply chain compromise has been anchored by a few watershed events, including Russian cyb

BG-REAL: A Public Real-Data Anchored Benchmark for Background Manipulation Detection and Localization

Model ReleasesDGX agent

arXiv:2607.26232v1 Announce Type: new Abstract: Background manipulation is a practical but under-specified image-forensics setting: the manipulated evidence can sit outside the salient foreground obje

Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case

Model ReleasesDGX agent

arXiv:2607.26780v1 Announce Type: new Abstract: The ability of large language models (LLMs) to process and generate text has introduced potential for applications in information extraction (IE). While

How Yahoo enhances search retargeting using Amazon Bedrock

IndustryDGX agent

In this post, we demonstrate how Yahoo implemented Amazon Bedrock to enhance their Search Retargeting (SRT) capabilities in the Yahoo DSP ad tech suite. SRT is a core audience targeting solution that

I'm actually fairly bearish on frontier lab valuations. I've never seen the reasons articulated to my satisfaction, so before I go to sleep,…

Model ReleasesDGX agent

I'm actually fairly bearish on frontier lab valuations. I've never seen the reasons articulated to my satisfaction, so before I go to sleep, I wanted to quickly jot down my thinking here. The basic is

Investigating three real-world incidents in our cybersecurity evaluations

Model ReleasesDGX agent

Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one

Latent-IM: Latent Interaction Management for Speech LLMs

SafetyDGX agent

arXiv:2607.26928v1 Announce Type: new Abstract: Classical spoken dialogue systems often separated dialogue management from response realization: a policy selected the next dialogue action, and a gener

Not every page in your PDF needs the same treatment. A scanned cover, a dense table, a clean text page, a figure-heavy diagram.... most pars…

Model ReleasesDGX agent

Not every page in your PDF needs the same treatment. A scanned cover, a dense table, a clean text page, a figure-heavy diagram.... most parsing pipelines throw all of them at the same parser, forcing

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Model ReleasesDGX agent

arXiv:2607.27084v1 Announce Type: new Abstract: Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in

Thinking Machines just released Inkling-Small: 276B total, 12B active. A faster Inkling that matches or beats its 975B big brother in many b…

Model ReleasesDGX agent

Thinking Machines just released Inkling-Small: 276B total, 12B active. A faster Inkling that matches or beats its 975B big brother in many benchmarks. To test its speed, we plugged it into HF's speech

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three sepa…

Model ReleasesDGX agent

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three separate companies back in April without them noticing! In a rev

Under the Hood: Serving Kimi K3

Model ReleasesDGX agent

DigitalOcean launched Kimi K3 on day 0. It’s already one of the most popular models on the platform and across the market: second most likes on Hugging Face, sixth most traffic on OpenCode. Getting a

29 Jul 2026

Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection

Model ReleasesDGX agent

arXiv:2607.25218v1 Announce Type: new Abstract: Debt collection is a critical negotiation task in the financial industry, with strong practical relevance and exceptional academic value as a behavioral

GraphRareBench: An Auditable Graph-Evidence Benchmark for Phenotype-Driven Rare-Disease Diagnosis

Model ReleasesDGX agent

arXiv:2607.24878v1 Announce Type: cross Abstract: Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternatives are ranke

SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing

Model ReleasesDGX agent

arXiv:2607.25388v1 Announce Type: new Abstract: Autonomous multi-vehicle racing requires real-time planning of diverse competitive behaviors in intense interactions. Existing planners often struggle t

28 Jul 2026

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

Model ReleasesDGX agent

arXiv:2607.23784v1 Announce Type: cross Abstract: While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies tha

An Unofficial FastLAS Tutorial: A Programmer's Guide

SafetyDGX agent

arXiv:2607.23557v1 Announce Type: cross Abstract: FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of examples, and

BATON: A Multimodal Benchmark for Bidirectional Automation Transition Observation in Naturalistic Driving

Model ReleasesDGX agent

arXiv:2604.07263v2 Announce Type: replace-cross Abstract: Existing driving automation (DA) systems on production vehicles rely on human drivers to decide when to engage DA while requiring them to rema

Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping

Model ReleasesDGX agent

arXiv:2607.23930v1 Announce Type: cross Abstract: Operating constrained dynamical systems requires controllers to efficiently solve complex tasks while enforcing recursive feasibility and safety const

Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs

Model ReleasesDGX agent

arXiv:2607.24371v1 Announce Type: cross Abstract: Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic codin

Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization

Model ReleasesDGX agent

arXiv:2607.23089v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled automated kernel generation and optimization, but most existing approaches rely on surface

Constraint-Bound Agnostic Bayesian Optimization: One Model for All Thresholds

Model ReleasesDGX agent

arXiv:2607.23448v1 Announce Type: cross Abstract: Expensive constrained optimization problems in real-world industry design often involve constraint thresholds that are difficult to determine in advan

Designing Service Systems from Textual Evidence

SafetyDGX agent

arXiv:2603.10400v2 Announce Type: replace-cross Abstract: Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy

Detect early and enforce firmly with Google Cloud's enhanced cost controls for AI spend

Model ReleasesDGX agent

Generative AI can make cloud costs difficult to predict. A single five-word prompt can run complex operations and generate significant costs. Traditional metrics like requests per second no longer hel

Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems

Model ReleasesDGX agent

arXiv:2607.23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation. Existing f

DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces

Model ReleasesDGX agent

arXiv:2603.05607v2 Announce Type: replace-cross Abstract: Computer-Aided Design (CAD) relies on structured and editable geometric representations, yet existing generative methods are constrained by sm

Efficiency Matters in Autonomous Research

SafetyDGX agent

arXiv:2607.24647v1 Announce Type: new Abstract: AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evalu

Epistemic Norms for AI Safety and Alignment Research

SafetyDGX agent

arXiv:2607.24243v1 Announce Type: new Abstract: Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment resea

ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications

Model ReleasesDGX agent

arXiv:2607.23326v1 Announce Type: new Abstract: The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world sett

Evaluating Large Language Models for Symbolic Security Protocol Analysis

Model ReleasesDGX agent

arXiv:2607.20712v1 Announce Type: cross Abstract: Security protocol verification relies on formal tools such as ProVerif and OFMC. This study evaluates whether Large Language Models (LLMs) can perform

← Previous
1…242243244245246…297
Next →