AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,724 results
Safety

Causal-JEPA: Learning World Models through Object-Level Latent Masking

DGX agent

arXiv:2602.11389v2 Announce Type: replace Abstract: World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a u

safetyarxiv-cs-ai
29 May 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk

DGX agent

arXiv:2605.29788v1 Announce Type: new Abstract: Critical sequential decisions are rarely single-timescale: a strategic decision causally shapes the context in which every subsequent tactical choice is

safetyarxiv-cs-ai
29 May 2026
Research

Conf-Gen: Conformal Uncertainty Quantification for Generative Models

DGX agent

arXiv:2605.28920v1 Announce Type: cross Abstract: Conformal prediction (CP) and its extension, conformal risk control (CRC), are established frameworks for quantifying uncertainty in supervised machin

researcharxiv-cs-ai
29 May 2026
Model Releases

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation

DGX agent

arXiv:2605.30000v1 Announce Type: new Abstract: Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

DiffSpot: Can VLMs Spot Fine-Grained Visual Differences in Web Interfaces?

DGX agent

arXiv:2605.29615v1 Announce Type: cross Abstract: Vision-language models (VLMs) have made strong progress on high-level image-text alignment, yet their ability to perceive subtle visual differences re

model-releasesarxiv-cs-cl
29 May 2026
Safety

Emergent Semantic Representations in World Models through Physical Interaction without Linguistic Supervision

DGX agent

arXiv:2605.28865v1 Announce Type: cross Abstract: What does a world model learn from physical exploration, without any linguistic supervision? We argue the answer is organized by a single principle: t

safetyarxiv-cs-ai
29 May 2026
Model Releases

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

DGX agent

arXiv:2605.28882v1 Announce Type: cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However,

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Hands-on with Gemini Spark beta rolling out to AI Ultra subs: planned a birthday party from emails and calendar, but called a live-in boyfriend a 'close friend' (Reece Rogers/Wired)

DGX agent

Reece Rogers / Wired: Hands-on with Gemini Spark beta rolling out to AI Ultra subs: planned a birthday party from emails and calendar, but called a live-in boyfriend a “close friend” — Google's new AI

model-releasestechmeme
29 May 2026
Tools

Here's everything you need to know about Replit in 60 seconds ⭐️ → Plain English prompts turned into real working software → End-to-end work…

DGX agent

Here's everything you need to know about Replit in 60 seconds ⭐️ → Plain English prompts turned into real working software → End-to-end workflow from UI to deployment → Real-time team collaboration wi

toolsreplit--x
29 May 2026
Safety

Jailbreaking and Mitigation of Vulnerabilities in Large Language Models

DGX agent

arXiv:2410.15236v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have transformed artificial intelligence by advancing natural language understanding and generation, enabling app

safetyarxiv-cs-ai
29 May 2026
Tutorials

MEMENTO: Leveraging Web as a Learning Signal for Low-Data Domains

DGX agent

arXiv:2605.29795v1 Announce Type: new Abstract: Real-world tasks often lack large labeled datasets, motivating extensive work on learning in low-data regimes. However, existing approaches such as few-

tutorialsarxiv-cs-ai
29 May 2026
Model Releases

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

DGX agent

arXiv:2605.29737v1 Announce Type: cross Abstract: LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code t

model-releasesarxiv-cs-cl
29 May 2026
Local Ai

New LFM2.5 8b A1b model!!

DGX agent

Liquid AI released LFM2.5-8B-A1B, a device-optimized model designed to power real-life applications on phones, laptops, PCs, robots, and lightweight server-side use-cases. The model is a fast, memory-

local-air-ollama
29 May 2026
Model Releases

OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation

DGX agent

arXiv:2605.29829v1 Announce Type: new Abstract: Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient par

model-releasesarxiv-cs-ai
29 May 2026
Safety

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

DGX agent

arXiv:2605.29582v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide pro

safetyarxiv-cs-cl
29 May 2026
Safety

REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image

DGX agent

arXiv:2605.30338v1 Announce Type: new Abstract: Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applic

safetyarxiv-cs-cv
29 May 2026
Hardware

Run Step 3.7 Flash on NVIDIA GPUs with Enterprise-Ready Multimodal AI

DGX agent

Step 3.7 Flash is a 198B-parameter Mixture-of-Experts vision-language model designed for enterprise-scale production workloads, featuring native image and video input, a 256k context window, and confi

hardwarenvidia-developer
29 May 2026
Research

SalsaAgent: A multimodal embodied language model for interactive dance generation

DGX agent

arXiv:2605.29219v1 Announce Type: new Abstract: Interaction between humanoids involves bidirectional and nonverbal reactivity, coordination and synchrony. Toward socially aware robots and interactive

researcharxiv-cs-cv
29 May 2026
Model Releases

The improvements run wide. Across all major European languages, Command A+ consistently pulls ahead of competitors on WMT24++ (xCOMET-XL): …

DGX agent

The improvements run wide. Across all major European languages, Command A+ consistently pulls ahead of competitors on WMT24++ (xCOMET-XL): 🇫🇷 +2.4 pts in French 🇪🇸 +1.9 pts in Spanish 🇩🇪 +0.9 pts in G

model-releasescohere--x
29 May 2026
Model Releases

The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More

DGX agent

arXiv:2603.23971v2 Announce Type: replace-cross Abstract: Developers and consumers increasingly choose reasoning models (RMs) based on their listed API prices. However, how accurately do these prices

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

When LLM Reward Design Fails: Diagnostic-Driven Refinement for Sparse Structured RL

DGX agent

arXiv:2605.28918v1 Announce Type: new Abstract: For sparse, structured reinforcement-learning tasks with semantic reward-function interfaces, LLM-generated reward shaping is better framed as debugging

model-releasesarxiv-cs-lg
29 May 2026
Safety

yes, absolutely, many companies are experimenting. but also: most of those experiments are failing to yield significant RoI. (weird for an e…

DGX agent

yes, absolutely, many companies are experimenting. but also: most of those experiments are failing to yield significant RoI. (weird for an economist to not even ask or address that question.) Looks li

safetygary-marcus--x
29 May 2026
Model Releases

AI researchers ran 15-day simulations of worlds governed by different AI models: Claude Sonnet 4.6 recorded no crimes, while Gemini 3 Flash had the most at 683 (Jake Angelo/Fortune)

DGX agent

Jake Angelo / Fortune: AI researchers ran 15-day simulations of worlds governed by different AI models: Claude Sonnet 4.6 recorded no crimes, while Gemini 3 Flash had the most at 683 — Imagine a world

model-releasestechmeme
28 May 2026
Research

Are We Truly Innovating? A Qualitative and Quantitative Study of Originality in AI Research Papers

DGX agent

arXiv:2602.06054v3 Announce Type: replace Abstract: Assessing originality in AI research is arguably the most consequential yet least reliable step in peer review. Reviewer judgments of originality re

researcharxiv-cs-cl
28 May 2026
Model Releases

Claude Opus 4.8: 'a modest but tangible improvement'

DGX agent

Anthropic shipped Claude Opus 4.8 today. My favourite thing about it is this note in the release announcement: Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor. Ther

model-releasessimon-willison
28 May 2026
Industry

Codex Mobile is free for everyone now and Goal Mode left beta. What it means if you do not code

DGX agent

Codex Mobile is now available in preview on iOS and Android across all ChatGPT plans, including Free and Go, in all supported regions. The feature allows developers to follow and steer Codex while it

industryr-chatgpt
28 May 2026
Safety

Commit to the Bit: Reactive Reinforcement Learning Done Right

DGX agent

arXiv:2605.28276v1 Announce Type: new Abstract: Reinforcement learning algorithms are commonly analyzed (and designed) under the Markov assumption. This is unrealistic, as most environments encountere

safetyarxiv-cs-lg
28 May 2026
Applications

Data Formulator 0.7: AI-powered data analytics for enterprise data

DGX agent

Data Formulator introduces AI-powered analytics for enterprise data workflows. Data teams can easily bring enterprise data into an AI-ready workspace where users can explore, analyze, and visualize da

applicationsmicrosoft-research
28 May 2026
Model Releases

Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models

DGX agent

arXiv:2512.00349v2 Announce Type: replace Abstract: Are frontier AI systems becoming more capable? Certainly. Yet such progress is not an unalloyed blessing but rather a Trojan horse: behind their per

model-releasesarxiv-cs-ai
28 May 2026
Safety

Delay-Aware Reinforcement Learning for Highway On-Ramp Merging under Stochastic Communication Latency

DGX agent

arXiv:2403.11852v5 Announce Type: replace-cross Abstract: Delayed and partially observable state information poses significant challenges for reinforcement learning (RL)-based control in real-world au

safetyarxiv-cs-ai
28 May 2026
Model Releases

DisasterBench: Benchmarking LLM Planning under Typed Tool Interface Constraints

DGX agent

arXiv:2605.27957v1 Announce Type: new Abstract: Disasters cause severe societal impacts, demanding rapid coordination of heterogeneous AI tools, from satellite analysis to flood prediction and damage

model-releasesarxiv-cs-cl
28 May 2026
Research

EchoAvatar: Real-time Generative Avatar Animation from Audio Streams

DGX agent

arXiv:2605.28272v1 Announce Type: new Abstract: Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistant

researcharxiv-cs-cv
28 May 2026
Model Releases

ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations

DGX agent

arXiv:2605.27908v1 Announce Type: cross Abstract: Existing emotional support conversation (ESC) systems mainly rely on end-to-end response generation or coarse strategy supervision, offering limited i

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Evolving Dataflow to process massive datasets for machine learning

DGX agent

Google created MapReduce more than 20 years ago to solve the scaling problems in data processing that the then young company was running into. The AI era that we are in now demands efficient, large-sc

model-releasesgoogle-cloud-ai
28 May 2026
Model Releases

How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving

DGX agent

arXiv:2605.28302v1 Announce Type: cross Abstract: Modern large language model (LLM) inference has progressively disaggregated to keep pace with growing model sizes and tight TTFT and TPOT service-leve

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

just noticed today - the dataset is already past 1k+ downloads. opensource / openresearch ftw ! @evo__hq would be opensourcing as many datas…

DGX agent

just noticed today - the dataset is already past 1k+ downloads. opensource / openresearch ftw ! @evo__hq would be opensourcing as many datasets, evals and autoresearch runs as we can in our pursuit of

model-releasesclem-delangue--x
28 May 2026
Model Releases

Local Reachy Mini conversations wireless looks like magic! You can bring your friend around the house and get WOW effect from anyone! 🔥 Tha…

DGX agent

Local Reachy Mini conversations wireless looks like magic! You can bring your friend around the house and get WOW effect from anyone! 🔥 Thanks @andimarafioti for the blog post on how to set this up: h

model-releasesclem-delangue--x
28 May 2026
Local Ai

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration

DGX agent

arXiv:2605.28805v1 Announce Type: cross Abstract: Visual outcomes are increasingly central to multimodal large language models, making reliable and fine-grained verification essential for scaling gene

local-aiarxiv-cs-ai
28 May 2026
Applications

OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models

DGX agent

arXiv:2605.27916v1 Announce Type: cross Abstract: The advancement of general medical Multimodal Large Language Models (MLLMs) has shown great potential for building conversational assistants to suppor

applicationsarxiv-cs-cl
28 May 2026
Model Releases

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

DGX agent

arXiv:2605.27378v1 Announce Type: new Abstract: Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have pro

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Pittsburgh-based Gray Swan, which stress-tests AI models for top frontier AI labs, raised a 40M Series A at a 200M valuation co-led by Wing VC and Madrona (Rashi Shrivastava/Forbes)

DGX agent

Rashi Shrivastava / Forbes: Pittsburgh-based Gray Swan, which stress-tests AI models for top frontier AI labs, raised a 40M Series A at a 200M valuation co-led by Wing VC and Madrona — Gray Swan works

model-releasestechmeme
28 May 2026
Model Releases

Plug-and-Play Benchmarking of Reinforcement Learning Algorithms for Large-Scale Flow Control

DGX agent

arXiv:2601.15015v2 Announce Type: replace Abstract: Reinforcement learning (RL) has shown promising results in active flow control (AFC), yet progress in the field remains difficult to assess as exist

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation

DGX agent

arXiv:2605.28237v1 Announce Type: cross Abstract: Real-world navigation is fundamentally driven by Points of Interest (POIs), yet reaching a precise POI remains a critical 'final-meters' challenge. Ex

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Snowveil: A Framework for Decentralised Preference Discovery

DGX agent

arXiv:2512.18444v2 Announce Type: replace-cross Abstract: Aggregating subjective preferences in social choice traditionally assumes a trusted central authority. In contrast, this paper formalises Dece

model-releasesarxiv-cs-ai
28 May 2026
Safety

Teacher-Student Representational Alignment for Reinforcement Learning-Driven Imitation Learning

DGX agent

arXiv:2605.28372v1 Announce Type: new Abstract: Imitation learning (IL) from a state-based reinforcement learning (RL) policy is a common approach to overcome the curse of dimensionality in complex an

safetyarxiv-cs-lg
28 May 2026
Safety

The Illusion of Opting in AI-Mediated Consequential Decisions

DGX agent

arXiv:2605.28210v1 Announce Type: new Abstract: Drawing on Ullmann-Margalit's concept of opting (transformative, irrevocable, and shadowed by foreclosed alternatives), we show that current AI systems

safetyarxiv-cs-ai
28 May 2026
Tutorials

“The mechanism is always the same in every story I've been covering. The demo works in a controlled environment with clean inputs. The deplo…

DGX agent

“The mechanism is always the same in every story I've been covering. The demo works in a controlled environment with clean inputs. The deployment fails because real kitchens, real intersections, and r

tutorialsgary-marcus--x
28 May 2026
Safety

VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking

DGX agent

arXiv:2605.28083v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models have emerged as powerful generalist policies, their severe vulnerability to adversarial patches significantly

safetyarxiv-cs-cv
28 May 2026
← Previous
1…345346347348349…370
Next →