AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
90,914Total entries
1Added by human
90,913Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
90,913 results
Tutorials

we just revamped the create_agent docs! the new agents page shows how to build a custom harness for your use case w/ create_agent as an easy…

DGX agent

we just revamped the create_agent docs! the new agents page shows how to build a custom harness for your use case w/ create_agent as an easy entrypoint start w/ your prompt and tools, then add middlew

tutorialsharrison-chase--x
26 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Tools

> … [W]e keep finding things that are mysterious, even unsettling. We find structures that mirror results from human neuroscience. We find e…

DGX agent

> … [W]e keep finding things that are mysterious, even unsettling. We find structures that mirror results from human neuroscience. We find evidence of introspection. We find internal states that funct

toolsboris-cherny--x
26 May 2026
Industry

We sat down with DeepMind CEO @demishassabis to find out how close we are to AGI, curing every disease with AI, and human meaning post-AGI. …

DGX agent

We sat down with DeepMind CEO @demishassabis to find out how close we are to AGI, curing every disease with AI, and human meaning post-AGI. Full interview below, and the rundown is coming to the newsl

industryrowan-cheung--x
26 May 2026
Tutorials

Weakly Supervised Camouflaged Object Detection Based on the SAM Model and Mask Guidance

DGX agent

arXiv:2605.25385v1 Announce Type: cross Abstract: Camouflaged object detection (COD) from a single image is a challenging task due to the high similarity between objects and their surroundings. Existi

tutorialsarxiv-cs-ai
26 May 2026
Industry

We're starting to see some PC makers respond to Apple's MacBook Neo

DGX agent

Apple's 599 MacBook Neo launched in March 2026 at approximately half the price of the MacBook Air, directly competing with lower-cost Windows PCs and triggering significant industry reaction. PC maker

industryars-technica
26 May 2026
Model Releases

We’ve shipped a security-guidance plugin for Claude Code that helps identify and fix vulnerabilities as you’re writing code. Available for a…

DGX agent

We’ve shipped a security-guidance plugin for Claude Code that helps identify and fix vulnerabilities as you’re writing code. Available for all Claude Code users. Install from the plugin marketplace (/

model-releasesboris-cherny--x
26 May 2026
Local Ai

What Are We Actually Decoding? Source Attribution for Non-Invasive Brain-to-Language Retrieval

DGX agent

arXiv:2605.24524v1 Announce Type: cross Abstract: In non-invasive neural language decoding, results can be inflated by sources that are not stimulus-evoked neural evidence: decoder priors, embedding-b

local-aiarxiv-cs-cl
26 May 2026
Safety

What Gets Cited: Competitive GEO in AI Answer Engines

DGX agent

arXiv:2605.25517v1 Announce Type: new Abstract: AI answer engines generate answers from retrieved pages but cite only a few sources. This makes visibility depend not just on ranking, but on being cite

safetyarxiv-cs-ai
26 May 2026
Applications

What Happens Next? Anticipating Future Motion by Generating Point Trajectories

DGX agent

arXiv:2509.21592v2 Announce Type: replace-cross Abstract: We consider the problem of forecasting motion from a single image, i.e., predicting how objects in the world are likely to move, without the a

applicationsarxiv-cs-ai
26 May 2026
Model Releases

What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA

DGX agent

arXiv:2605.25988v1 Announce Type: new Abstract: Medical RAG needs evidence-grounded claims, so plugging a claim-level NLI checker into retrieval-augmented RL is intuitive. extbf{We find that the check

model-releasesarxiv-cs-cl
26 May 2026
Research

What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics

DGX agent

arXiv:2510.16435v2 Announce Type: replace-cross Abstract: With the growing use of large language models and conversational interfaces in human-robot interaction, robots' ability to answer user questio

researcharxiv-cs-cl
26 May 2026
Industry

What to expect at FinOps X 2026: Join theCUBE June 8-10

DGX agent

Financial operations, known as FinOps, are undergoing a quiet shift — from managing cloud costs to shaping business decisions based on AI-driven insights. The practice is now an essential tool for man

industrysiliconangle
26 May 2026
Model Releases

When Can We Trust Early Warnings? Leakage-Excluded Early Outcome Prediction from LMS Interaction Logs

DGX agent

arXiv:2605.25794v1 Announce Type: new Abstract: Early-warning models built from Learning Management System (LMS) logs aim to predict end-of-course outcomes early enough to enable timely learner suppor

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure

DGX agent

arXiv:2605.23932v1 Announce Type: new Abstract: Despite strong medical benchmark accuracy, LLMs can exhibit severe multi-turn sycophancy in clinical dialogue, abandoning initial correct diagnosis unde

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

DGX agent

arXiv:2605.25981v1 Announce Type: new Abstract: We document an empirical phenomenon in chain-of-thought and ReAct agents driven by ten large language models from seven architecture families: meaning-b

model-releasesarxiv-cs-cl
26 May 2026
Safety

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

DGX agent

arXiv:2605.24202v1 Announce Type: new Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learn

safetyarxiv-cs-ai
26 May 2026
Research

When Does Synthetic Patent Data Help? Volume-Fidelity Trade-offs in Low-Resource Multi-Label Classification

DGX agent

arXiv:2605.24296v1 Announce Type: new Abstract: We study when LLM-generated synthetic data helps low-resource multi-label patent classification, separating true synthetic value from the confound that

researcharxiv-cs-ai
26 May 2026
Research

When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges

DGX agent

arXiv:2605.26046v1 Announce Type: cross Abstract: Customizing an LLM judge to a specific task or domain often involves optimizing its prompt across multiple evaluation criteria simultaneously. Textual

researcharxiv-cs-ai
26 May 2026
Tools

When I woke up this morning I didn't think I'd be spending a bunch of time today getting familiar with Catholic theology, but here we are. N…

DGX agent

When I woke up this morning I didn't think I'd be spending a bunch of time today getting familiar with Catholic theology, but here we are. Notes on Pope Leo XIV's encyclical on AI. https://simonwillis

toolssimon-willison--x
26 May 2026
Research

When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift

DGX agent

arXiv:2605.25629v1 Announce Type: new Abstract: Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train--t

researcharxiv-cs-cl
26 May 2026
Research

When Interpretability Becomes a Liability: Adversarial Attacks on CBM Concept Layers

DGX agent

arXiv:2605.25304v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) have emerged as a cornerstone approach for interpretable machine learning, providing human-understandable intermediate

researcharxiv-cs-lg
26 May 2026
Model Releases

When Mean CE Fails: Median CE Can Better Track Language Model Quality

DGX agent

arXiv:2605.24667v1 Announce Type: new Abstract: Mean cross-entropy is the standard validation metric for language models, but it can fail to track model quality during training. We examine this in two

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation

DGX agent

arXiv:2605.24902v1 Announce Type: cross Abstract: Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured clinical do

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When Search Becomes Memory: Turning Robot Design Trials into Transferable Skills

DGX agent

arXiv:2605.25832v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as proposal generators for evolutionary robot design, yet most loops remain memoryless: simulator r

model-releasesarxiv-cs-ai
26 May 2026
Safety

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

DGX agent

arXiv:2605.25864v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewar

safetyarxiv-cs-cl
26 May 2026
Model Releases

When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents

DGX agent

arXiv:2605.24069v1 Announce Type: cross Abstract: The rise of tool-using Large Language Model (LLM) agents, standardized by protocols like the Model Context Protocol (MCP), has unlocked unprecedented

model-releasesarxiv-cs-ai
26 May 2026
Safety

when you think about it SpaceX has cumulatively lost a lot less than OpenAI so maybe it’s a bargain? or maybe we just shouldn’t value compan…

DGX agent

when you think about it SpaceX has cumulatively lost a lot less than OpenAI so maybe it’s a bargain? or maybe we just shouldn’t value companies at a trillion dollars until they actually show evidence

safetygary-marcus--x
26 May 2026
Model Releases

WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems

DGX agent

arXiv:2605.24579v1 Announce Type: new Abstract: Long-context memory systems often fail under fixed budgets, but end-to-end evaluation does not reveal whether evidence was discarded during compression

model-releasesarxiv-cs-cl
26 May 2026
Research

Which Is Better For Reducing Outdated and Vulnerable Dependencies: Pinning or Floating?

DGX agent

arXiv:2510.08609v3 Announce Type: replace-cross Abstract: Developers consistently use version constraints to specify acceptable versions of the dependencies for their project. Pinning dependencies can

researcharxiv-cs-lg
26 May 2026
Applications

WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers

DGX agent

arXiv:2509.10452v2 Announce Type: replace Abstract: Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance. In man

applicationsarxiv-cs-cl
26 May 2026
Industry

who did it better

DGX agent

This post likely features a comparison of two similar images, designs, or creations side-by-side asking viewers to judge which version is superior—a common format on social media for crowdsourcing opi

industryemad-mostaque--x
26 May 2026
Model Releases

Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring

DGX agent

arXiv:2605.24737v1 Announce Type: cross Abstract: Current approaches to AI compliance treat conformity as a binary, audit-time verdict rather than a continuous, measurable property of production syste

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

WhoSaidIt: Human-LLM Collaborative Annotation for Text-Based Multilingual Speaker-Attribute Classification

DGX agent

arXiv:2605.26070v1 Announce Type: new Abstract: Annotating speaker attributes from text is inherently ambiguous, particularly in multilingual settings where demographic and social cues are implicit an

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

DGX agent

arXiv:2605.25256v1 Announce Type: new Abstract: Aligning AI systems with organizational decision-making is typically framed as a single-target problem: make the model behave like the organization. We

model-releasesarxiv-cs-ai
26 May 2026
Agents

Why Agentic Theorem Prover Works: A Statistical Provability Theory of Mathematical Reasoning Models

DGX agent

arXiv:2602.10538v3 Announce Type: replace-cross Abstract: Agentic theorem provers combine a reasoning model, retrieval, search, and a proof assistant verifier, yet it remains unclear which components

agentsarxiv-cs-lg
26 May 2026
Agents

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform

DGX agent

arXiv:2605.23972v1 Announce Type: new Abstract: Large language models achieve strong performance in language generation and knowledge-intensive tasks, yet remain limited in settings requiring causal r

agentsarxiv-cs-ai
26 May 2026
Agents

Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

DGX agent

arXiv:2601.22984v2 Announce Type: replace Abstract: Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evalua

agentsarxiv-cs-ai
26 May 2026
Model Releases

WideDepth: Millimeter-Accurate Benchmark for Fisheye Depth Estimation

DGX agent

arXiv:2605.24074v1 Announce Type: cross Abstract: Fisheye cameras are increasingly adopted in robotics for near-field manipulation, navigation, and immersive perception, yet indoor depth benchmarks wi

model-releasesarxiv-cs-ro
26 May 2026
Applications

Windows' classic 3D Space Cadet pinball is getting a physical re-creation

DGX agent

The Windows 95 classic 3D Space Cadet pinball game is being converted from its original digital format into a functional physical pinball machine. The recreation uses largely 3D-printed components inc

applicationsars-technica
26 May 2026
Research

WINO: A Weak-Form Physics Informed Neural Operator for Hyperelasticity on Variable Domains

DGX agent

arXiv:2605.24651v1 Announce Type: cross Abstract: We propose a Weak-form Physics-Informed Neural Operator (WINO), a data-free framework that combines the efficiency of neural operators with the geomet

researcharxiv-cs-lg
26 May 2026
Applications

WISE: Web Information Satire and Fakeness Evaluation

DGX agent

arXiv:2512.24000v3 Announce Type: replace Abstract: Distinguishing fake or untrue news from satire or humor poses a unique challenge due to their overlapping linguistic features and divergent intent.

applicationsarxiv-cs-cl
26 May 2026
Model Releases

WLNO: Wavelet-Laplace Neural Operator for Solving Partial Differential Equations

DGX agent

arXiv:2605.24658v1 Announce Type: new Abstract: This work introduces the Wavelet-Laplace Neural Operator (WLNO), a novel neural operator that fuses Haar wavelet multi-scale spatial decomposition with

model-releasesarxiv-cs-lg
26 May 2026
Research

Word Class Representations Spontaneously Emerge from Successor Representations Trained on Natural Language

DGX agent

arXiv:2605.24585v1 Announce Type: new Abstract: Language models are typically trained to predict the next token in a sequence. Here, we explore an alternative predictive principle from reinforcement l

researcharxiv-cs-cl
26 May 2026
Safety

Workflow cleanup tools for ComfyUI: Visual Fold, group folding, and node alignment

DGX agent

Visual Fold is a tool for simple visual organization of ComfyUI workflows that does not turn selected nodes into a subgraph or change workflow logic. Group folding and node alignment features enable c

safetyr-stablediffusion
26 May 2026
Model Releases

World-State Transformations for Neuro-symbolic Interactive Storytelling

DGX agent

arXiv:2605.24719v1 Announce Type: cross Abstract: Large Language Models (LLMs) have changed the possibilities of Interactive Storytelling systems that process free-text user input. However, as more of

model-releasesarxiv-cs-ai
26 May 2026
Safety

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy

DGX agent

arXiv:2602.06508v2 Announce Type: replace Abstract: Reinforcement learning (RL) can refine Vision-Language-Action (VLA) policies beyond behavior cloning, but real-world RL remains expensive due to ext

safetyarxiv-cs-ro
26 May 2026
Model Releases

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point

DGX agent

arXiv:2502.08047v5 Announce Type: replace Abstract: Recent progress in GUI agents has substantially improved visual grounding, yet robust planning remains challenging, particularly when the environmen

model-releasesarxiv-cs-ai
26 May 2026
Research

WTKO-CNN: Deep Learning Reveals Sequence Motifs Distinguishing Wild-Type and Knockout ATAC-seq Peaks

DGX agent

arXiv:2605.24034v1 Announce Type: cross Abstract: Chromatin regulators can alter transcriptional programs by modifying the accessibility of regulatory DNA elements. Understanding how regulatory sequen

researcharxiv-cs-ai
26 May 2026
← Previous
1…10891090109110921093…1895
Next →