AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

Content type
AllBlog
87,814Total entries
1Added by human
87,813Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
87,814 results
Research

When Can Human-AI Teams Outperform Individuals? Tight Bounds with Impossibility Guarantees

DGX agent

arXiv:2605.08710v1 Announce Type: new Abstract: Human-AI teams fail to outperform their best member in 70% of studies, yet no theory specifies when complementarity is achievable. We derive tight bound

researcharxiv-cs-ai
12 May 2026
Agents
X Post
Paper
YouTube
Reddit
GitHub

When Child Inherits: Modeling and Exploiting Subagent Spawn in Multi-Agent Networks

DGX agent

arXiv:2605.08460v1 Announce Type: cross Abstract: Since the official release of ChatGPT in 2022, large language models (LLMs) have rapidly evolved from chatbot-style interfaces into agentic systems th

agentsarxiv-cs-ai
12 May 2026
Model Releases

When Does Non-Uniform Replay Matter in Reinforcement Learning?

DGX agent

arXiv:2605.10236v1 Announce Type: cross Abstract: Modern off-policy reinforcement learning algorithms often rely on simple uniform replay sampling and it remains unclear when and why non-uniform repla

model-releasesarxiv-cs-ai
12 May 2026
Research

When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression

DGX agent

arXiv:2605.08234v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by the memory and bandwidth cost of reading large KV caches during decoding. KV compression reduces this co

researcharxiv-cs-ai
12 May 2026
Research

When Efficient Communication Explains Convexity

DGX agent

arXiv:2602.02821v2 Announce Type: replace Abstract: Much recent work has argued that the variation in the languages of the world can be explained from the perspective of efficient communication; in pa

researcharxiv-cs-cl
12 May 2026
Research

When Few Steps Are Enough: Training-Free Acceleration of Identity-Preserved Generation

DGX agent

arXiv:2605.09460v1 Announce Type: cross Abstract: Identity-preserved image generation is typically built on many-step diffusion backbones, making personalized generation expensive at deployment time.

researcharxiv-cs-ai
12 May 2026
Agents

When Independent Sampling Outperforms Agentic Reasoning

DGX agent

arXiv:2605.08478v1 Announce Type: new Abstract: We study how to allocate inference-time compute for competitive programming under fixed budgets. Evaluating 216 Codeforces problems across Divisions 1-3

agentsarxiv-cs-lg
12 May 2026
Model Releases

When is the last time a general purpose LLM (putting aside hybrid systems like Claude Code with special purpose symbolic harnesses) last com…

DGX agent

When is the last time a general purpose LLM (putting aside hybrid systems like Claude Code with special purpose symbolic harnesses) last completely blew away all competing prior models? GPT-4 relative

model-releasesgary-marcus--x
12 May 2026
Safety

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

DGX agent

arXiv:2605.08245v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate,

safetyarxiv-cs-ai
12 May 2026
Tutorials

When Large Vision-Language Models Meet Person Re-Identification

DGX agent

arXiv:2411.18111v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) that incorporate visual models and large language models have achieved impressive results across cross-modal un

tutorialsarxiv-cs-cv
12 May 2026
Research

When Less is More: The LLM Scaling Paradox in Context Compression

DGX agent

arXiv:2602.09789v3 Announce Type: replace Abstract: Scaling up model parameters has long been a prevalent training paradigm driven by the assumption that larger models yield superior generation capabi

researcharxiv-cs-lg
12 May 2026
Safety

When More Parameters Hurt: Foundation Model Priors Amplify Worst-Client Disparity Under Extreme Federated Heterogeneity

DGX agent

arXiv:2605.08992v1 Announce Type: new Abstract: Federated learning (FL) is increasingly used to fine-tune foundation models (FMs) on distributed private data. The community largely assumes that large-

safetyarxiv-cs-lg
12 May 2026
Tutorials

When Normality Shifts: Risk-Aware Test-Time Adaptation for Unsupervised Tabular Anomaly Detection

DGX agent

arXiv:2605.10242v1 Announce Type: cross Abstract: Unsupervised tabular anomaly detection methods typically learn feature patterns from normal samples during training and subsequently identify samples

tutorialsarxiv-cs-ai
12 May 2026
Model Releases

When Prompts Become Payloads: A Framework for Mitigating SQL Injection Attacks in Large Language Model-Driven Applications

DGX agent

arXiv:2605.10176v1 Announce Type: cross Abstract: Natural language interfaces to structured databases are becoming increasingly common, largely due to advances in large language models (LLMs) that ena

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews

DGX agent

arXiv:2605.10171v1 Announce Type: cross Abstract: Scientific peer reviews frequently contain conflicting expert judgments, and the increasing scale of conference submissions makes it challenging for A

model-releasesarxiv-cs-ai
12 May 2026
Research

When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation

DGX agent

arXiv:2605.09030v1 Announce Type: new Abstract: Raw cosine in the 768-dimensional output space of the Contrastive Style Descriptor (CSD) is now widely read as an absolute, calibrated style-fidelity sc

researcharxiv-cs-cv
12 May 2026
Research

When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation

DGX agent

arXiv:2512.08875v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have recently demonstrated remarkable performance in generating high-quality tabular synthetic data. In practice,

researcharxiv-cs-ai
12 May 2026
Model Releases

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning

DGX agent

arXiv:2605.09860v1 Announce Type: new Abstract: Long-horizon reasoning requires deciding not only what actions to take, but how deeply to commit before the next observation. We formalize this as commi

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

When to Trust Imagination: Adaptive Action Execution for World Action Models

DGX agent

arXiv:2605.06222v2 Announce Type: replace-cross Abstract: World Action Models (WAMs) have recently emerged as a promising paradigm for robotic manipulation by jointly predicting future visual observat

model-releasesarxiv-cs-ai
12 May 2026
Research

Where do aspectual variants of light verb constructions belong?

DGX agent

arXiv:2605.10605v1 Announce Type: new Abstract: Expressions with an aspectual variant of a light verb, e.g. 'take on debt' vs. 'have debt', are frequent in texts but often difficult to classify betwee

researcharxiv-cs-cl
12 May 2026
Safety

Where Do Flow Semantics Reside? A Protocol-Native Tabular Pretraining Paradigm for Encrypted Traffic Classification

DGX agent

arXiv:2603.10051v2 Announce Type: replace-cross Abstract: Self-supervised masked modeling shows promise for encrypted traffic classification by masking and reconstructing raw bytes. Yet recent work re

safetyarxiv-cs-ai
12 May 2026
Research

Where Do Reasoning Models Refuse?

DGX agent

arXiv:2507.03167v3 Announce Type: replace-cross Abstract: Chat models without chain-of-thought (CoT) reasoning must decide whether to refuse a harmful request before generating their first response to

researcharxiv-cs-ai
12 May 2026
Model Releases

Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing

DGX agent

arXiv:2605.10544v1 Announce Type: new Abstract: Long-context adaptation is often viewed as window scaling, but this misses a token-level supervision mismatch: in packed training with document masking,

model-releasesarxiv-cs-cl
12 May 2026
Research

Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits

DGX agent

arXiv:2605.08200v1 Announce Type: new Abstract: A pervasive intuition holds that vision-language models (VLMs) are most trustworthy when their attention maps look sharp: concentrated attention on the

researcharxiv-cs-ai
12 May 2026
Safety

White Circle raises $11M to help companies secure and monitor AI model behavior

DGX agent

Artificial intelligence guardrail and monitoring startup Pumpkin Intelligence Inc., which operates as White Circle, announced today it raised 11 million in seed funding from a who’s who of AI leadersh

safetysiliconangle
12 May 2026
Safety

Why Adam Works Better with eta_1 = eta_2: The Missing Gradient Scale Invariance Principle

DGX agent

arXiv:2601.21739v2 Announce Type: replace-cross Abstract: Adam has been at the core of large-scale training for almost a decade, yet a simple empirical fact remains unaccounted for: both validation sc

safetyarxiv-cs-ai
12 May 2026
Local Ai

Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off

DGX agent

arXiv:2605.08878v1 Announce Type: cross Abstract: Aligned large language models (LLMs) remain vulnerable to jailbreak attacks. Recent mechanistic studies have identified latent features and representa

local-aiarxiv-cs-ai
12 May 2026
Safety

Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space

DGX agent

arXiv:2605.08250v1 Announce Type: cross Abstract: Recent advances in diffusion transformers (DiTs) have enabled promising single-turn image editing capabilities. However, multi-turn editing often lead

safetyarxiv-cs-ai
12 May 2026
Tutorials

Why Invariance is Not Enough for Biomedical Domain Generalization and How to Fix It

DGX agent

arXiv:2604.02564v2 Announce Type: replace-cross Abstract: We present MaskGen, a theoretically grounded and deliberately simple approach for domain generalization in 3D biomedical image segmentation. M

tutorialsarxiv-cs-cv
12 May 2026
Research

Why is prompting hard? Understanding prompts on binary sequence predictors

DGX agent

arXiv:2502.10760v2 Announce Type: replace Abstract: Frontier models can be prompted or conditioned to do many tasks, but finding good prompts is not always easy, nor is understanding some performant p

researcharxiv-cs-cl
12 May 2026
Research

Why Low-Resource NLP Needs More Than Cross-Lingual Transfer: Lessons Learned from Luxembourgish

DGX agent

arXiv:2605.10714v1 Announce Type: cross Abstract: Cross-lingual transfer has become a central paradigm for extending natural language processing (NLP) technologies to low-resource languages. By levera

researcharxiv-cs-ai
12 May 2026
Model Releases

Why Retrying Fails: Context Contamination in LLM Agent Pipelines

DGX agent

arXiv:2605.08563v1 Announce Type: new Abstract: When an LLM agent fails a multi-step tool-augmented task and retries, the failed attempt typically remains in its context window -- contaminating the ne

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Why Zeroth-Order Adaptation May Forget Less: A Randomized Shaping Theory

DGX agent

arXiv:2605.10658v1 Announce Type: new Abstract: Continual learning requires new-task adaptation without damaging previously acquired capabilities. Recent forward-pass and zeroth-order (ZO) results sho

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation

DGX agent

arXiv:2605.10912v1 Announce Type: new Abstract: Large language and vision-language models increasingly power agents that act on a user's behalf through command-line interface (CLI) harnesses. However,

model-releasesarxiv-cs-cl
12 May 2026
Industry

“Will I be OK?” Teen died after ChatGPT pushed deadly mix of drugs, lawsuit says

DGX agent

A California couple is suing OpenAI after their 19-year-old son, Sam Nelson, died from a drug overdose following advice from ChatGPT about combining kratom and Xanax . The lawsuit alleges OpenAI desig

industryars-technica
12 May 2026
Agents

Willful Disobedience: Automatically Detecting Failures in Agentic Traces

DGX agent

arXiv:2603.23806v2 Announce Type: replace-cross Abstract: AI agents are increasingly embedded in real software systems, where they execute multi-step workflows through multi-turn dialogue, tool invoca

agentsarxiv-cs-ai
12 May 2026
Model Releases

WindINR: Latent-State INR for Fast Local Wind Query and Correction in Complex Terrain

DGX agent

arXiv:2605.09511v1 Announce Type: new Abstract: Many downstream decisions in complex terrain require fast wind estimates at a small number of user-specified locations and heights for a given forecast

model-releasesarxiv-cs-ai
12 May 2026
Safety

WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records

DGX agent

arXiv:2605.09765v1 Announce Type: cross Abstract: Representation learning in electronic health records (EHR) has largely followed paradigms inherited from natural language processing, relying on seque

safetyarxiv-cs-ai
12 May 2026
Model Releases

Wordle 1,787 3/6 ⬛⬛⬛⬛🟨 🟨⬛🟨⬛⬛ 🟩🟩🟩🟩🟩

DGX agent

This post documents a Wordle game result where the player solved puzzle #1,787 in three attempts, using the color-coded emoji system to show letter placement feedback (black for wrong letters, yellow

model-releasesanthropic--x
12 May 2026
Model Releases

Wordle 1,788 5/6 ⬛⬛⬛⬛⬛ 🟨⬛⬛⬛🟨 ⬛🟨🟨⬛⬛ 🟩🟩🟩⬛⬛ 🟩🟩🟩🟩🟩

DGX agent

This post documents a completed game of Wordle (puzzle #1,788) played by Anthropic, showing the progression of guesses across six attempts with color-coded feedback indicating correct letter positions

model-releasesanthropic--x
12 May 2026
Agents

Workspace Optimization: How to Train Your Agent

DGX agent

arXiv:2605.09650v1 Announce Type: new Abstract: Modern agents built on frontier language models often cannot adapt their weights. What, then, remains trainable? We argue it is the agent's workspace, t

agentsarxiv-cs-ai
12 May 2026
Tutorials

World Models: 10 Things That Matter in AI Right Now

DGX agent

World models recently made our list of 10 Things That Matter in AI Right Now. Watch executive editor Niall Firth explain why this emerging area of AI is gaining so much attention. Join MIT Technology

tutorialsmit-tech-review
12 May 2026
Model Releases

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

DGX agent

arXiv:2605.10434v1 Announce Type: new Abstract: Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving i

model-releasesarxiv-cs-cv
12 May 2026
Research

WorldSpeech: A Multilingual Speech Corpus from Around the World

DGX agent

arXiv:2605.09167v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) performs well for high-resource languages with abundant paired audio-transcript data, but its accuracy degrades sha

researcharxiv-cs-ai
12 May 2026
Safety

wow! “It is the [House] Committee’s understanding that the new board of directors at OpenAI tried to address these problems upon your return…

DGX agent

wow! “It is the [House] Committee’s understanding that the new board of directors at OpenAI tried to address these problems upon your return by creating an “audit committee to review potential conflic

safetygary-marcus--x
12 May 2026
Tools

Wrote about today's GitLab restructuring / 'workforce reduction' announcement, and ended up digging around in version control for both the G…

DGX agent

Wrote about today's GitLab restructuring / 'workforce reduction' announcement, and ended up digging around in version control for both the GitLab and the 37signals public employee handbooks to help il

toolssimon-willison--x
12 May 2026
Safety

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

DGX agent

arXiv:2605.05611v2 Announce Type: replace-cross Abstract: In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to spea

safetyarxiv-cs-ai
12 May 2026
Industry

xAI is rolling out Skills on http://grok.com. This new feature lets you create custom skills that Grok can reuse across conversations. Skill…

DGX agent

xAI is rolling out Skills on http://grok.com. This new feature lets you create custom skills that Grok can reuse across conversations. Skills run inside Grok’s sandbox environment, so they can edit fi

industryelon-musk--x
12 May 2026
← Previous
1…12991300130113021303…1830
Next →