AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,608 results
Safety

Reliability-Oriented Multilingual Orthopedic Diagnosis: A Domain-Adaptive Modeling and a Conceptual Validation Framework

DGX agent

arXiv:2605.02266v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly proposed for clinical decision support including multilingual diagnosis in low-resource settings. However,

safetyarxiv-cs-cl
5 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Hardware

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving

DGX agent

arXiv:2605.01708v1 Announce Type: cross Abstract: Contemporary systems serving large language models (LLMs) have adopted prefill-decode disaggregation to better load-balance between the compute-bound

hardwarearxiv-cs-lg
5 May 2026
Safety

STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

DGX agent

arXiv:2605.02122v1 Announce Type: new Abstract: Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fr

safetyarxiv-cs-lg
5 May 2026
Model Releases

STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Storie

DGX agent

arXiv:2601.08510v3 Announce Type: replace Abstract: Movie screenplays are rich long-form narratives that interleave complex character relationships, temporally ordered events, and dialogue-driven inte

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation

DGX agent

arXiv:2508.15658v5 Announce Type: replace Abstract: The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show pr

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

toodles from mickey mouse clubhouse was weirdly ahead of its time wake phrase: mickey mouse clubhouse launched in 2006, and “oh toodles” tra…

DGX agent

toodles from mickey mouse clubhouse was weirdly ahead of its time wake phrase: mickey mouse clubhouse launched in 2006, and “oh toodles” trained toddlers on the assistant wake phrase years before siri

model-releasesyohei-nakajima--x
5 May 2026
Research

Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments

DGX agent

arXiv:2505.09901v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate or automate human behavior in complex sequential decision-making settings. A na

researcharxiv-cs-cl
4 May 2026
Safety

Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations

DGX agent

arXiv:2512.20260v5 Announce Type: replace Abstract: Weakly-Supervised Camouflaged Object Detection (WSCOD) aims to locate and segment objects that are visually concealed within their surrounding scene

safetyarxiv-cs-cv
4 May 2026
Safety

Decentralized Proximal Stochastic Gradient Langevin Dynamics

DGX agent

arXiv:2605.00723v1 Announce Type: cross Abstract: We propose Decentralized Proximal Stochastic Gradient Langevin Dynamics (DE-PSGLD), a decentralized Markov chain Monte Carlo (MCMC) algorithm for samp

safetyarxiv-cs-lg
4 May 2026
Model Releases

FollowTable: A Benchmark for Instruction-Following Table Retrieval

DGX agent

arXiv:2605.00400v1 Announce Type: cross Abstract: Table Retrieval (TR) has traditionally been formulated as an ad-hoc retrieval problem, where relevance is primarily determined by topical semantic sim

model-releasesarxiv-cs-cl
4 May 2026
Safety

Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding

DGX agent

arXiv:2605.00642v1 Announce Type: cross Abstract: Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capabili

safetyarxiv-cs-cv
4 May 2026
Safety

Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values

DGX agent

arXiv:2605.00762v1 Announce Type: new Abstract: We propose a new framework for meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback (BCMAB-FBF). Unlike semi-ba

safetyarxiv-cs-lg
4 May 2026
Model Releases

Minimizing Human Intervention in Online Classification

DGX agent

arXiv:2510.23557v2 Announce Type: replace-cross Abstract: Training or fine-tuning large language model (LLM)-based systems often requires costly human feedback, yet there is limited understanding of h

model-releasesarxiv-cs-lg
4 May 2026
Safety

Optimizing Resource-Constrained Non-Pharmaceutical Interventions for Multi-Cluster Outbreak Control Using Hierarchical Reinforcement Learning

DGX agent

arXiv:2603.19397v2 Announce Type: replace Abstract: Non-pharmaceutical interventions (NPIs), such as diagnostic testing and quarantine, are crucial for controlling infectious disease outbreaks but are

safetyarxiv-cs-lg
4 May 2026
Model Releases

Redis Array Playground

DGX agent

Tool: Redis Array Playground Salvatore Sanfilippo submitted a PR adding a new data type - arrays - to Redis. The new commands are ARCOUNT, ARDEL, ARDELRANGE, ARGET, ARGETRANGE, ARGREP, ARINFO, ARINSER

model-releasessimon-willison
4 May 2026
Safety

Reinforcement Learning with LLM-Guided Action Spaces for Synthesizable Lead Optimization

DGX agent

arXiv:2604.07669v2 Announce Type: replace Abstract: Lead optimization in drug discovery requires improving therapeutic properties while ensuring that molecular modifications correspond to feasible syn

safetyarxiv-cs-lg
4 May 2026
Safety

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

DGX agent

arXiv:2605.00380v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diver

safetyarxiv-cs-cl
4 May 2026
Model Releases

Semantic Level of Detail for Knowledge Graphs: Discovering Abstraction Boundaries via Spectral Heat Diffusion

DGX agent

arXiv:2603.08965v2 Announce Type: replace Abstract: Graph-structured knowledge systems -- from knowledge graphs to GraphRAG pipelines -- organize information into hierarchical communities, yet lack a

model-releasesarxiv-cs-lg
4 May 2026
Applications

This is why I keep going on about LangGraph 1.2 alpha is looking great: node-level error handlers, timeouts, improved performance, and stron…

DGX agent

This is why I keep going on about LangGraph 1.2 alpha is looking great: node-level error handlers, timeouts, improved performance, and stronger typing LangGraph is the runtime layer you want underneat

applicationsharrison-chase--x
4 May 2026
Tools

this one is doing v well btw if you want the popular vote filter on the firehose of all the things @patrickdebois was one of the track keyno…

DGX agent

this one is doing v well btw if you want the popular vote filter on the firehose of all the things @patrickdebois was one of the track keynotes i gave a 'blank check' to based on his sincere support s

toolsswyx--x
4 May 2026
Safety

World Model for Robot Learning: A Comprehensive Survey

DGX agent

arXiv:2605.00080v1 Announce Type: cross Abstract: World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They s

safetyarxiv-cs-cv
4 May 2026
Safety

Are AIs about to subjugate humanity? In the debate about catastrophic AI risk, evolutionary scenarios receive far too little attention compa…

DGX agent

Are AIs about to subjugate humanity? In the debate about catastrophic AI risk, evolutionary scenarios receive far too little attention compared to largely speculative arguments about “instrumental con

safetydan-hendrycks--x
2 May 2026
Model Releases

Grok Voice is used by Starlink right now

DGX agent

Grok Voice is used by Starlink right now Grok Voice brutally dominates the top of the τ-voice Bench Grok scores 67.3%, while Gemini sits at 43.8% and GPT Realtime at 35.3% This is a massive lead over

model-releaseselon-musk--x
2 May 2026
Applications

I was quoted a couple times in this Atlantic article, but that isn’t (the only) reason I think it is good. It lays out the reasons why we wh…

DGX agent

I was quoted a couple times in this Atlantic article, but that isn’t (the only) reason I think it is good. It lays out the reasons why we whipsawed from “AI is a bubble” to “there are not enough data

applicationsethan-mollick--x
2 May 2026
Safety

People are really enjoying our full workshops showing end to end walkthroughs of real production workflows! This is a rare double header wit…

DGX agent

People are really enjoying our full workshops showing end to end walkthroughs of real production workflows! This is a rare double header with @braintrust's Giran Moodley and @OussamaHaff walking thoug

safetyswyx--x
2 May 2026
Industry

Apple may take 'several months' to catch up to Mac mini and Studio demand

DGX agent

During its Q2 2026 earnings call, Apple CEO Tim Cook stated that Mac mini and Mac Studio may take several months to reach supply-demand balance. Apple underestimated demand for these products, which h

industryars-technica
1 May 2026
Model Releases

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR

DGX agent

arXiv:2604.27543v1 Announce Type: new Abstract: Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into sh

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

ChipLingo: A Systematic Training Framework for Large Language Models in EDA

DGX agent

arXiv:2604.27415v1 Announce Type: new Abstract: With the rapid advancement of semiconductor technology, Electronic Design Automation (EDA) has become an increasingly knowledge-intensive and document-d

model-releasesarxiv-cs-lg
1 May 2026
Safety

Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations

DGX agent

arXiv:2604.27372v1 Announce Type: cross Abstract: This paper investigates the continuous-time counterpart of the Q-function for entropy-regularized mean-field control (MFC) with controlled common nois

safetyarxiv-cs-lg
1 May 2026
Safety

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning…

DGX agent

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning fixes get bolted on at post-training. By then, the patterns

safetydair-ai--x
1 May 2026
Model Releases

From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction

DGX agent

arXiv:2604.27906v1 Announce Type: new Abstract: Persistent AI memory is often reduced to a retrieval problem: store prior interactions as text, embed them, and ask the model to recover relevant contex

model-releasesarxiv-cs-ai
1 May 2026
Applications

IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development

DGX agent

arXiv:2604.16399v2 Announce Type: replace-cross Abstract: The widespread adoption of AI-assisted development tools in 2025 -- and the emergence of vibe coding, a practice of generating complete applic

applicationsarxiv-cs-ai
1 May 2026
Model Releases

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions

DGX agent

arXiv:2604.27763v1 Announce Type: new Abstract: The emergence of Large Language Models (LLMs) offers a transformative interface for Web3, yet existing benchmarks fail to capture the complexity of tran

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

KellyBench: A Benchmark for Long-Horizon Sequential Decision Making

DGX agent

arXiv:2604.27865v1 Announce Type: new Abstract: Language models are saturating benchmarks for procedural tasks with narrow objectives. But they are increasingly being deployed in long-horizon, non-sta

model-releasesarxiv-cs-ai
1 May 2026
Safety

Knowledge Graph Representations for LLM-Based Policy Compliance Reasoning

DGX agent

arXiv:2604.27713v1 Announce Type: new Abstract: The risks posed by AI features are increasing as they are rapidly integrated into software applications. In response, regulations and standards for safe

safetyarxiv-cs-ai
1 May 2026
Model Releases

Language Models Refine Mechanical Linkage Designs Through Symbolic Reflection and Modular Optimisation

DGX agent

arXiv:2604.27962v1 Announce Type: new Abstract: Designing mechanical linkages involves combinatorial topology selection and continuous parameter fitting. We show that language models can systematicall

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

On 5/5 @realDanFu and team will discuss DSV4’s hybrid attention and KV cache efficiency, should be a great session!

DGX agent

On 5/5 @realDanFu and team will discuss DSV4’s hybrid attention and KV cache efficiency, should be a great session! Join us Tue 5/5: #DeepSeek-V4's hybrid attention + sparse MoE reduces KV cache up to

model-releasestogether-ai--x
1 May 2026
Model Releases

One week since the launch of GPT-5.5, and it’s already our strongest model launch yet. API revenue is growing more than 2x faster than any p…

DGX agent

One week since the launch of GPT-5.5, and it’s already our strongest model launch yet. API revenue is growing more than 2x faster than any prior release, while Codex doubled revenue in under seven day

model-releasesopenai--x
1 May 2026
Model Releases

Our CEO @jerryjliu0 in @VentureBeat , on what's actually changing in the LLM stack: 'We've really identified that there's a core set of data…

DGX agent

Our CEO @jerryjliu0 in @VentureBeat , on what's actually changing in the LLM stack: 'We've really identified that there's a core set of data that has been locked up in all these file format containers

model-releasesjerry-liu--x
1 May 2026
Hardware

Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference

DGX agent

arXiv:2604.26968v1 Announce Type: cross Abstract: Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference serving. Current

hardwarearxiv-cs-ai
1 May 2026
Model Releases

The new Grok comes in below the latest Chinese open weights models, Grok 4 was at the frontier when released. (& Artificial Analysis: please…

DGX agent

The new Grok comes in below the latest Chinese open weights models, Grok 4 was at the frontier when released. (& Artificial Analysis: please stop using GDPval-AA which is not a useful test of anything

model-releasesethan-mollick--x
1 May 2026
Model Releases

TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering

DGX agent

arXiv:2604.28076v1 Announce Type: cross Abstract: Large Language Models (LLMs) have advanced Table Question Answering, where most queries can be answered by extracting information or simple aggregatio

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

When training Grok 4.3, we spoke directly with devs and businesses to understand what they actually needed: a model that’s fast, affordable,…

DGX agent

When training Grok 4.3, we spoke directly with devs and businesses to understand what they actually needed: a model that’s fast, affordable, and great at tool calling. The result is a daily driver tha

model-releaseselon-musk--x
1 May 2026
Safety

A Scaled Three-Vehicle Platooning Platform

DGX agent

arXiv:2604.25963v1 Announce Type: new Abstract: Vehicle platooning has attracted increasing attention as a promising approach to improve traffic efficiency, energy consumption, and roadway safety thro

safetyarxiv-cs-ro
30 Apr 2026
Safety

A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

DGX agent

arXiv:2510.08049v3 Announce Type: replace-cross Abstract: Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward m

safetyarxiv-cs-ai
30 Apr 2026
Applications

AGI is GPT-X doing all of the work after you say 'We would love you to throw a party for yourself as a marketing event for OpenAI, so do tha…

DGX agent

This post by Ethan Mollick likely discusses how advanced AI systems like hypothetical future GPT versions could autonomously execute complex, real-world tasks with minimal human direction, using a hum

applicationsethan-mollick--x
30 Apr 2026
Hardware

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving

DGX agent

arXiv:2604.26103v1 Announce Type: cross Abstract: All current LLM serving systems place the GPU at the center, from production-level attention-FFN disaggregation to NVIDIA's Rubin GPU-LPU heterogeneou

hardwarearxiv-cs-ai
30 Apr 2026
Applications

Appian puts reliability at the center of enterprise AI as accuracy gaps frustrate organizations

DGX agent

Enterprise software has no shortage of AI enthusiasm, but the gap between AI potential and production-ready results continues to frustrate organizations investing in enterprise process automation. As

applicationssiliconangle
30 Apr 2026
← Previous
1…352353354355356…367
Next →