AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,608 results
11 May 2026

Cloud Storage Rapid: Turbocharged object storage for AI and analytics

Model ReleasesDGX agent

At Google Cloud Next ’26 we announced Cloud Storage Rapid, a family of object storage capabilities for data-intensive workloads like AI and analytics. Out of the gate, Cloud Storage Rapid consists of

DeepSeek V4 Flash is ~90% cheaper than GPT 5.4 Mini and ~70% cheaper than Gemini 3.1 Flash Lite For devs pushing ~500M tok/month, this is th…

Model ReleasesDGX agent

DeepSeek V4 Flash is ~90% cheaper than GPT 5.4 Mini and ~70% cheaper than Gemini 3.1 Flash Lite For devs pushing ~500M tok/month, this is the difference between: GPT 5.4 Mini: ~394/mo Gemini 3.1 Flash

Do Joint Audio-Video Generation Models Understand Physics?

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.07061v1 Announce Type: cross Abstract: Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visu

EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams

Model ReleasesDGX agent

arXiv:2605.07299v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users

Emergent social transmission of model-based representations without inference

SafetyDGX agent

arXiv:2604.05777v2 Announce Type: replace Abstract: How do people acquire rich, flexible knowledge about their environment from others despite limited cognitive capacity? Humans are often thought to r

How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study

Model ReleasesDGX agent

arXiv:2605.05340v2 Announce Type: replace-cross Abstract: As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awa

InvThink: Premortem Reasoning for Safer Language Models

SafetyDGX agent

arXiv:2510.01569v3 Announce Type: replace Abstract: We present InvThink, a training and prompting framework that requires the model to enumerate, analyze, and constrain potential failures before gener

Meet the latest Database Center, now with Gemini-powered fleet intelligence

Model ReleasesDGX agent

Managing a modern database fleet is both a scale and cognitive problem. As database estates grow, the effort required to monitor, troubleshoot, and optimize them often outpaces teams’ capacity, who fi

MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants

Model ReleasesDGX agent

arXiv:2603.09652v3 Announce Type: replace Abstract: With the rapid advancement of Large Language Models (LLMs) in code generation, human-AI interaction is evolving from static text responses to dynami

Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models

SafetyDGX agent

arXiv:2605.07649v1 Announce Type: cross Abstract: Over the last few years, research on autonomous systems has matured to such a degree that the field is increasingly well-positioned to translate resea

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation

Model ReleasesDGX agent

arXiv:2605.07334v1 Announce Type: new Abstract: Video Reasoning Segmentation (VRS) aims to segment target objects in videos based on implicit instructions that convey human intent and temporal logic.

Revisiting Adam for Streaming Reinforcement Learning

ResearchDGX agent

arXiv:2605.06764v1 Announce Type: cross Abstract: Learning from a sequence of interactions, as soon as observations are perceived and acted upon, without explicitly storing them, holds the promise of

RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation

Model ReleasesDGX agent

arXiv:2605.07129v1 Announce Type: cross Abstract: Large Language Models (LLMs) have emerged as a promising paradigm for next-generation recommender systems, offering strong semantic understanding and

S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models

Model ReleasesDGX agent

arXiv:2503.05085v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have fundamentally reshaped speech-to-speech (S2S) systems, enabling increasingly natural spoken int

Semantic State Abstraction Interfaces for LLM-Augmented Portfolio Decisions: Multi-Axis News Decomposition and RL Diagnostics

ResearchDGX agent

arXiv:2605.06730v1 Announce Type: new Abstract: We introduce Semantic State Abstraction Interfaces (SSAI): a methodological template for mapping sparse unstructured text into K auditable, named coordi

🧵 Slime: The Most Elegant & Comfortable RL Training Framework Ever A deep dive into why Slime redefines LLM RL training with clean architec…

HardwareDGX agent

🧵 Slime: The Most Elegant & Comfortable RL Training Framework Ever A deep dive into why Slime redefines LLM RL training with clean architecture & production-grade engineering ✨ Insights from Zhihu con

The Position Curse: LLMs Struggle to Locate the Last Few Items in a List

Model ReleasesDGX agent

arXiv:2605.07127v1 Announce Type: cross Abstract: Modern large language models (LLMs) can find a needle in a haystack (locating a single relevant fact buried among hundreds of thousands of irrelevant

Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles

Local AiDGX agent

arXiv:2512.03454v4 Announce Type: replace-cross Abstract: Interpreting natural-language commands to localize target objects is critical for autonomous driving (AD). Existing visual grounding (VG) meth

“we gonna yolo our way into running the biggest conf in town.” and somehow… we actually did it with @aiDotEngineer singapore a year ago, @ag…

ToolsDGX agent

“we gonna yolo our way into running the biggest conf in town.” and somehow… we actually did it with @aiDotEngineer singapore a year ago, @agrimsingh @unprofeshme and i joked about the idea of running

Your AI Use Is Breaking My Brain

TutorialsDGX agent

Your AI Use Is Breaking My Brain Excellent, angry piece by Jason Koebler on how AI writing online is becoming impossible to avoid, filtering it is mentally exhausting and it's even starting to distort

10 May 2026

A look at Janitor AI, a romantic fantasy roleplay chatbot site run by three men that claims 2.5M DAUs and 15M total users, with 70% to 80% identifying as women (Anna Tong/Forbes)

ApplicationsDGX agent

Anna Tong / Forbes: A look at Janitor AI, a romantic fantasy roleplay chatbot site run by three men that claims 2.5M DAUs and 15M total users, with 70% to 80% identifying as women — Despite its name,

9 May 2026

100% agree on the Context Hub. Developers constantly tweak their approaches to manage context in their prompts with each new model and tool …

Model ReleasesDGX agent

100% agree on the Context Hub. Developers constantly tweak their approaches to manage context in their prompts with each new model and tool suite release. I would even say that the core problem we fac

ERNIE 5.1 is here 🚀 ERNIE 5.1 significantly reduces pretraining cost while compressing total parameters to ~1/3 and activated parameters to…

Model ReleasesDGX agent

ERNIE 5.1 is here 🚀 ERNIE 5.1 significantly reduces pretraining cost while compressing total parameters to ~1/3 and activated parameters to ~1/2 — using only ~6% of the pretraining cost compared to mo

I recently installed Hermes after raw-dogging Claude Code for a year or so. Have tried it at work, didn't see the need to implement it perso…

Model ReleasesDGX agent

I recently installed Hermes after raw-dogging Claude Code for a year or so. Have tried it at work, didn't see the need to implement it personally. Always had the usual fears about 'privacy' and what n

8 May 2026

Improving Bash Generation in Small Language Models with Grammar-Constrained Decoding

HardwareDGX agent

Grammar-constrained decoding modifies language model generation by applying grammar constraints at each step to block structurally invalid tokens , ensuring syntactically correct Bash command generati

Remember that MIT study that showed that the ROI for generative AI wasn’t really there for most businesses? Or any of the six or seven studi…

SafetyDGX agent

Remember that MIT study that showed that the ROI for generative AI wasn’t really there for most businesses? Or any of the six or seven studies from other teams that followed, showing basically the sam

Thank you to @robertwiblin for inviting me on the @80000Hours podcast to discuss the research progress we’re making at @LawZero_ to create s…

TutorialsDGX agent

Thank you to @robertwiblin for inviting me on the @80000Hours podcast to discuss the research progress we’re making at @LawZero_ to create safe-by-design AI systems. Our current approach, Scientist AI

Which Macs are suffering from shortages—and where are things getting worse?

IndustryDGX agent

Apple CEO Tim Cook acknowledged in Q2 2026 earnings that high-demand Mac mini and Mac Studio configurations are severely constrained and may take several months to reach supply-demand balance. Apple h

7 May 2026

AsymmetryZero: A Framework for Operationalizing Human Expert Preferences as Semantic Evals

Model ReleasesDGX agent

arXiv:2605.04083v1 Announce Type: new Abstract: Much of the focus in RL today is on evaluation design: building meaningful evals that serve simultaneously as benchmarks and as well-defined reward sign

Congrats on the launch! Filesystems are all you need (?) There wasn't a huge demand for 'managed RAG' services in 2023, but it's possible th…

ApplicationsDGX agent

Congrats on the launch! Filesystems are all you need (?) There wasn't a huge demand for 'managed RAG' services in 2023, but it's possible the infra and market was just not mature enough. Maybe filesys

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems

Local AiDGX agent

arXiv:2605.03900v1 Announce Type: new Abstract: Frontier AI systems perform best in settings with clear, stable, and verifiable objectives, such as code generation, mathematical reasoning, games, and

Hacker News → LLM Artifact I built the most personalized HN feed. It only tracks topics I do research around based on memory and LLM wiki. N…

ResearchDGX agent

Hacker News → LLM Artifact I built the most personalized HN feed. It only tracks topics I do research around based on memory and LLM wiki. No point in storing bookmarks. With a few automations, rules,

I really enjoyed chatting with @mattturck, was a great discussion.

SafetyDGX agent

I really enjoyed chatting with @mattturck, was a great discussion. Deeply thoughtful conversation with @zicokolter, board member at @OpenAI and head of the machine learning department at @CarnegieMell

InterFuserDVS: Event-Enhanced Sensor Fusion for Safe RL-Based Decision Making

SafetyDGX agent

arXiv:2605.04355v1 Announce Type: new Abstract: Autonomous driving systems rely heavily on robust sensor fusion to perceive complex envi- ronments. Traditional setups using RGB cameras and LiDAR often

KGLAMP: Knowledge Graph-guided Language model for Adaptive Multi-robot Planning and Replanning

Model ReleasesDGX agent

arXiv:2602.04129v2 Announce Type: replace Abstract: Heterogeneous multi-robot systems are increasingly used in long-horizon missions requiring coordinated planning across diverse capabilities. However

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation

Local AiDGX agent

arXiv:2512.23864v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown remarkable generalization by mapping web-scale knowledge to robotic control, yet they remain bl

MenuNet: A Strategy-Proof Mechanism for Matching Markets

SafetyDGX agent

arXiv:2605.03216v1 Announce Type: cross Abstract: Strategy-proofness is a fundamental desideratum in mechanism design, ensuring truthful reporting and robust participation. Stability is another centra

PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World

ResearchDGX agent

arXiv:2605.05163v1 Announce Type: new Abstract: Synthesizing physics-grounded 3D assets is a critical bottleneck for interactive virtual worlds and embodied AI. Existing methods predominantly focus on

SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States

ResearchDGX agent

arXiv:2605.04496v1 Announce Type: new Abstract: Long-Text Understanding (LTU) at million-token scale requires balancing reasoning fidelity with computational efficiency. Frontier long-context LLMs can

Searching for the right context in a sea of information is such a timeless problem in our field. I'm excited to talk about some of the appro…

ToolsDGX agent

Searching for the right context in a sea of information is such a timeless problem in our field. I'm excited to talk about some of the approaches we've learned and benefited from at Thrive next week a

To Fuse or to Drop? Dual-Path Learning for Resolving Modality Conflicts in Multimodal Emotion Recognition

SafetyDGX agent

arXiv:2605.04877v1 Announce Type: cross Abstract: Multimodal emotion recognition (MER) benefits from combining text, audio, and vision, yet standard fusion often fails when modalities conflict. Crucia

We already had gemini-3.1-flash-lite-preview back on March 3rd, not clear if this new gemini-3.1-flash-lite is different other than no longe…

Model ReleasesDGX agent

We already had gemini-3.1-flash-lite-preview back on March 3rd, not clear if this new gemini-3.1-flash-lite is different other than no longer being marked as a 'preview'. Pricing appears to be the sam

We've teamed up with @cerebras to offer free Windsurf plans for SWE-1.6 Fast Mode at up to 1000 tok/s! Fast Mode is built on Cerebras infere…

Model ReleasesDGX agent

We've teamed up with @cerebras to offer free Windsurf plans for SWE-1.6 Fast Mode at up to 1000 tok/s! Fast Mode is built on Cerebras inference, enabling superior speed for planning and development wi

6 May 2026

An explainable hypothesis-driven approach to Drug-Induced Liver Injury with HADES

Model ReleasesDGX agent

arXiv:2605.02669v1 Announce Type: new Abstract: Drug-induced liver injury (DILI) remains a leading cause of late-stage clinical trial attrition. However, existing computational predictors primarily re

Audio-Visual Intelligence in Large Foundation Models

SafetyDGX agent

arXiv:2605.04045v1 Announce Type: new Abstract: Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines

Causal Software Engineering: A Vision and Roadmap

Model ReleasesDGX agent

arXiv:2605.02454v1 Announce Type: cross Abstract: Software engineering increasingly involves making high-stakes decisions under uncertainty, using signals from code, field data, and socio-technical pr

DiagramNet: An End-to-End Recognition Framework and Dataset for Non-Standard System-Level Diagrams

Model ReleasesDGX agent

arXiv:2605.01338v1 Announce Type: new Abstract: System-level diagrams encode the architectural blueprint of chip design, specifying module functions, dataflows, and interface protocols. However, non-s

False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models

SafetyDGX agent

arXiv:2601.07885v2 Announce Type: replace-cross Abstract: Emoticons are widely used in digital communication to convey affective intent, yet their safety implications for Large Language Models (LLMs)

HAAS: A Policy-Aware Framework for Adaptive Task Allocation Between Humans and Artificial Intelligence Systems

Model ReleasesDGX agent

arXiv:2605.02832v1 Announce Type: new Abstract: Deciding how to distribute work between humans and AI systems is a central challenge in organisational design. Most approaches treat this as a binary ch

Live blog: Code w/ Claude 2026

Model ReleasesDGX agent

Simon Willison's live blog covers Anthropic's Code w/ Claude 2026 event, documenting the morning keynote sessions with real-time updates. The event featured announcements including updates to Claude m

Model Spec Midtraining: Improving How Alignment Training Generalizes

SafetyDGX agent

arXiv:2605.02087v1 Announce Type: new Abstract: Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard a

On Verbalized Confidence Scores for LLMs

Model ReleasesDGX agent

arXiv:2412.14737v2 Announce Type: replace Abstract: The rise of large language models (LLMs) and their tight integration into our daily life make it essential to dedicate efforts towards their trustwo

OpenAI phone leaks: a push toward AI first hardware.

IndustryDGX agent

OpenAI is fast-tracking development of its first AI-focused smartphone with mass production targeted for early 2027 , marking the company's entry into consumer hardware . The device, positioned as an

PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization

Model ReleasesDGX agent

arXiv:2605.03129v1 Announce Type: cross Abstract: Browsing-enabled LLM assistants can fetch webpages and answer contact-seeking queries, creating a practical channel for scraping contact-style persona

@PPLXfinance @PPLXDevs On FinSearchComp T1, Finance Search delivered the highest accuracy for live financial data and the lowest cost per co…

ApplicationsDGX agent

@PPLXfinance @PPLXDevs On FinSearchComp T1, Finance Search delivered the highest accuracy for live financial data and the lowest cost per correct answer in the cohort. Every result includes citations,

this deepagents deploy https://docs.langchain.com/oss/python/deepagents/deploy (or at least directionally where we want to take it) what's m…

Model ReleasesDGX agent

this deepagents deploy https://docs.langchain.com/oss/python/deepagents/deploy (or at least directionally where we want to take it) what's missing? give us feedback! can someone PLEASE launch OS claud

Two weeks after release, Hy3 preview is #1 on @OpenRouter's weekly leaderboard with 3.66T tokens processed, up 298% week-over-week. #1 in ov…

Model ReleasesDGX agent

Two weeks after release, Hy3 preview is #1 on @OpenRouter's weekly leaderboard with 3.66T tokens processed, up 298% week-over-week. #1 in overall usage, tool calls, and coding. 15.4% market share acro

Valley3: Scaling Omni Foundation Models for E-commerce

Model ReleasesDGX agent

arXiv:2605.01278v1 Announce Type: new Abstract: In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understandi

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development

Model ReleasesDGX agent

arXiv:2603.04601v2 Announce Type: replace-cross Abstract: Code generation has emerged as one of AI's highest-impact use cases, yet existing benchmarks measure isolated tasks rather than the complete '

We built a simple app that's also probably the best PDF -> text phone app there. Take a picture of any document: a filled out form, identifi…

Local AiDGX agent

We built a simple app that's also probably the best PDF -> text phone app there. Take a picture of any document: a filled out form, identification, a statement, an essay - and we'll convert it into we

← Previous
1…280281282283284…294
Next →