AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
Model Releases

AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching

DGX agent

arXiv:2607.27130v1 Announce Type: new Abstract: Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only

model-releasesarxiv-cs-ai
31 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

DGX agent

arXiv:2607.26639v1 Announce Type: cross Abstract: A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% def

safetyarxiv-cs-ai
31 Jul 2026
Model Releases

Bridging AI and Energy Forecasting: An Autonomous Workflow with Customized Toolkit

DGX agent

arXiv:2307.07191v3 Announce Type: replace Abstract: Energy forecasting is crucial for the power grid, but fundamentally different from general time series analysis: it highly relies on covariates like

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

DGX agent

arXiv:2607.27747v1 Announce Type: new Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either

model-releasesarxiv-cs-cl
31 Jul 2026
Local Ai

Can Vision-Language Models Reason about AI Edits in Images?

DGX agent

arXiv:2607.28464v1 Announce Type: new Abstract: Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly

local-aiarxiv-cs-cv
31 Jul 2026
Safety

Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy

DGX agent

arXiv:2607.27212v1 Announce Type: cross Abstract: Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of q

safetyarxiv-cs-cl
31 Jul 2026
Model Releases

DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

DGX agent

arXiv:2607.27263v1 Announce Type: new Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

DGX agent

arXiv:2607.26611v1 Announce Type: new Abstract: AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in us

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

DGX agent

arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of do

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

DGX agent

Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without

model-releasesr-localllama
31 Jul 2026
Model Releases

LLMs struggle to simulate human belief updates in controlled environments

DGX agent

arXiv:2607.28347v1 Announce Type: new Abstract: LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been

model-releasesarxiv-cs-cl
31 Jul 2026
Hardware

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

DGX agent

arXiv:2607.28312v1 Announce Type: new Abstract: Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches prima

hardwarearxiv-cs-cv
31 Jul 2026
Model Releases

Optimal Realistic Local AI for Most

DGX agent

So you’ve got a 3090 or maybe even a 5090? Or more likely a 4060 8GB Ti. You wanna try local AI, you don’t know what it can/can’t do. 1) Install the best model you can. If you have a 3090 or a 5090, t

model-releasesr-localllama
31 Jul 2026
Model Releases

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

DGX agent

arXiv:2607.28008v1 Announce Type: new Abstract: Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthe

model-releasesarxiv-cs-cl
31 Jul 2026
Tutorials

RLPF: Reinforcement Learning from Performance Feedback for Code Generation

DGX agent

arXiv:2607.27271v1 Announce Type: new Abstract: Code models are increasingly trained with execution feedback, but most training signals still stop at correctness. This leaves an important gap for syst

tutorialsarxiv-cs-lg
31 Jul 2026
Model Releases

Scaling medical imaging report generation with multimodal reinforcement learning

DGX agent

arXiv:2601.17151v2 Announce Type: replace-cross Abstract: Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit ma

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

DGX agent

arXiv:2607.27421v1 Announce Type: new Abstract: Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable

model-releasesarxiv-cs-cl
31 Jul 2026
Safety

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

DGX agent

arXiv:2607.26566v1 Announce Type: cross Abstract: Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them inter

safetyarxiv-cs-ai
31 Jul 2026
Hardware

Simulation of Surgical Suturing Using Position-Based Dynamics and the Material Point Method for Robot Reinforcement Learning

DGX agent

arXiv:2607.27494v1 Announce Type: new Abstract: Recent advances in robotics research have created a strong demand for high-performance simulators. Surgical robotics simulation faces unique challenges

hardwarearxiv-cs-ro
31 Jul 2026
Model Releases

smevals - a small eval suite for evaluating models, prompts, and harnesses

DGX agent

smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer

model-releasessimon-willison
31 Jul 2026
Model Releases

THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model

DGX agent

arXiv:2607.27303v1 Announce Type: cross Abstract: Temporal heterogeneous graphs offer a natural abstraction for dynamic relational systems in which diverse node and relation types co-exist and evolve

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

DGX agent

arXiv:2607.28287v1 Announce Type: cross Abstract: ARC-AGI-3 turns abstraction into an interactive problem of skill acquisition. A player must infer an unfamiliar game's rules, hidden state, and goal w

model-releasesarxiv-cs-cv
31 Jul 2026
Model Releases

Very interesting paper on recursive self-improvement. The whole stack is released. Machine learning engineering gives recursive self-improve…

DGX agent

Very interesting paper on recursive self-improvement. The whole stack is released. Machine learning engineering gives recursive self-improvement a concrete, executable testbed. OpenMLE is an open full

model-releasesdair-ai--x
31 Jul 2026
Model Releases

APEX-Accounting

DGX agent

arXiv:2607.27189v1 Announce Type: new Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountant

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

ARC-Encoder: learning compressed text representations for large language models

DGX agent

arXiv:2510.20535v2 Announce Type: replace Abstract: Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs. Co

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Batten Down Your Packages: Mitigation Guidance for Supply Chain Compromise

DGX agent

Written by: Kelli Vanderlee, Stuart Carrera For years, the cybersecurity industry's understanding of software supply chain compromise has been anchored by a few watershed events, including Russian cyb

model-releasesgoogle-cloud-ai
30 Jul 2026
Model Releases

BG-REAL: A Public Real-Data Anchored Benchmark for Background Manipulation Detection and Localization

DGX agent

arXiv:2607.26232v1 Announce Type: new Abstract: Background manipulation is a practical but under-specified image-forensics setting: the manipulated evidence can sit outside the salient foreground obje

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case

DGX agent

arXiv:2607.26780v1 Announce Type: new Abstract: The ability of large language models (LLMs) to process and generate text has introduced potential for applications in information extraction (IE). While

model-releasesarxiv-cs-cl
30 Jul 2026
Industry

How Yahoo enhances search retargeting using Amazon Bedrock

DGX agent

In this post, we demonstrate how Yahoo implemented Amazon Bedrock to enhance their Search Retargeting (SRT) capabilities in the Yahoo DSP ad tech suite. SRT is a core audience targeting solution that

industryaws-ml-blog
30 Jul 2026
Model Releases

I'm actually fairly bearish on frontier lab valuations. I've never seen the reasons articulated to my satisfaction, so before I go to sleep,…

DGX agent

I'm actually fairly bearish on frontier lab valuations. I've never seen the reasons articulated to my satisfaction, so before I go to sleep, I wanted to quickly jot down my thinking here. The basic is

model-releasesgary-marcus--x
30 Jul 2026
Model Releases

Investigating three real-world incidents in our cybersecurity evaluations

DGX agent

Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one

model-releasessimon-willison
30 Jul 2026
Safety

Latent-IM: Latent Interaction Management for Speech LLMs

DGX agent

arXiv:2607.26928v1 Announce Type: new Abstract: Classical spoken dialogue systems often separated dialogue management from response realization: a policy selected the next dialogue action, and a gener

safetyarxiv-cs-cl
30 Jul 2026
Model Releases

Not every page in your PDF needs the same treatment. A scanned cover, a dense table, a clean text page, a figure-heavy diagram.... most pars…

DGX agent

Not every page in your PDF needs the same treatment. A scanned cover, a dense table, a clean text page, a figure-heavy diagram.... most parsing pipelines throw all of them at the same parser, forcing

model-releasesjerry-liu--x
30 Jul 2026
Model Releases

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

DGX agent

arXiv:2607.27084v1 Announce Type: new Abstract: Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

Thinking Machines just released Inkling-Small: 276B total, 12B active. A faster Inkling that matches or beats its 975B big brother in many b…

DGX agent

Thinking Machines just released Inkling-Small: 276B total, 12B active. A faster Inkling that matches or beats its 975B big brother in many benchmarks. To test its speed, we plugged it into HF's speech

model-releasessoumith-chintala--x
30 Jul 2026
Model Releases

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three sepa…

DGX agent

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three separate companies back in April without them noticing! In a rev

model-releasessimon-willison--x
30 Jul 2026
Model Releases

Under the Hood: Serving Kimi K3

DGX agent

DigitalOcean launched Kimi K3 on day 0. It’s already one of the most popular models on the platform and across the market: second most likes on Hugging Face, sixth most traffic on OpenCode. Getting a

model-releasesdigitalocean
30 Jul 2026
Model Releases

Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection

DGX agent

arXiv:2607.25218v1 Announce Type: new Abstract: Debt collection is a critical negotiation task in the financial industry, with strong practical relevance and exceptional academic value as a behavioral

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

GraphRareBench: An Auditable Graph-Evidence Benchmark for Phenotype-Driven Rare-Disease Diagnosis

DGX agent

arXiv:2607.24878v1 Announce Type: cross Abstract: Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternatives are ranke

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing

DGX agent

arXiv:2607.25388v1 Announce Type: new Abstract: Autonomous multi-vehicle racing requires real-time planning of diverse competitive behaviors in intense interactions. Existing planners often struggle t

model-releasesarxiv-cs-ro
29 Jul 2026
Model Releases

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

DGX agent

arXiv:2607.23784v1 Announce Type: cross Abstract: While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies tha

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

An Unofficial FastLAS Tutorial: A Programmer's Guide

DGX agent

arXiv:2607.23557v1 Announce Type: cross Abstract: FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of examples, and

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

BATON: A Multimodal Benchmark for Bidirectional Automation Transition Observation in Naturalistic Driving

DGX agent

arXiv:2604.07263v2 Announce Type: replace-cross Abstract: Existing driving automation (DA) systems on production vehicles rely on human drivers to decide when to engage DA while requiring them to rema

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping

DGX agent

arXiv:2607.23930v1 Announce Type: cross Abstract: Operating constrained dynamical systems requires controllers to efficiently solve complex tasks while enforcing recursive feasibility and safety const

model-releasesarxiv-cs-ro
28 Jul 2026
Model Releases

Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs

DGX agent

arXiv:2607.24371v1 Announce Type: cross Abstract: Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic codin

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization

DGX agent

arXiv:2607.23089v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled automated kernel generation and optimization, but most existing approaches rely on surface

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Constraint-Bound Agnostic Bayesian Optimization: One Model for All Thresholds

DGX agent

arXiv:2607.23448v1 Announce Type: cross Abstract: Expensive constrained optimization problems in real-world industry design often involve constraint thresholds that are difficult to determine in advan

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Designing Service Systems from Textual Evidence

DGX agent

arXiv:2603.10400v2 Announce Type: replace-cross Abstract: Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy

safetyarxiv-cs-ai
28 Jul 2026
← Previous
1…303304305306307…371
Next →