AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,603 results
8 Jun 2026

Towards Tight Bounds for Streaming Attention

Model ReleasesDGX agent

arXiv:2606.07205v1 Announce Type: cross Abstract: The attention mechanism is a cornerstone of modern transformer architectures. However, its expressive power comes at the cost of quadratic runtime and

Trading Engagement for Sustainability: Carbon-Aware Re-ranking for E-commerce Recommendations

Model ReleasesDGX agent

arXiv:2606.04550v1 Announce Type: cross Abstract: E-commerce recommender systems strongly influence which products users consider and purchase, yet sustainability signals such as Product Carbon Footpr

Tree-of-Experience: A Structured Experience-Management Solution for Self-Evolving Agents under Low-Repetition and Implicit-Reward Environments

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.06960v1 Announce Type: new Abstract: Experience-based self-evolution is crucial for LLM agents, but existing benchmarks often assume explicit goals, stable task patterns, and clear feedback

TSAQA: Time Series Analysis Question And Answering Benchmark

Model ReleasesDGX agent

arXiv:2601.23204v2 Announce Type: replace Abstract: Time series data are integral to critical applications across domains such as finance, healthcare, transportation, and environmental science. While

Twin: Tuning Learning Rate and Weight Decay of Deep Homogeneous Classifiers without Validation

Model ReleasesDGX agent

arXiv:2403.05532v2 Announce Type: replace-cross Abstract: We introduce Tune without Validation (Twin), a simple and effective pipeline for tuning learning rate and weight decay of homogeneous classifi

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.06673v1 Announce Type: new Abstract: Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak g

Uncertainty-Guided Label Rebalancing for CPS Safety Monitoring

Model ReleasesDGX agent

arXiv:2603.25670v3 Announce Type: replace Abstract: Safety monitoring is essential for Cyber-Physical Systems (CPSs). However, unsafe events are rare in real-world CPS operations, creating an extreme

Uniform Stability and Generalization Error of GD and SGD on Fixed-Point Parameters

Model ReleasesDGX agent

arXiv:2606.06934v1 Announce Type: new Abstract: We analyze generalization error, uniform stability, and uniform argument stability of gradient descent (GD) and stochastic gradient descent (SGD) over d

UniSHARP: Universal Sharp Monocular View Synthesis

Model ReleasesDGX agent

arXiv:2606.07514v1 Announce Type: new Abstract: In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across a continuum of cam

UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

Model ReleasesDGX agent

arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are

Unsupervised Continual Clustering via Forward-Backward Knowledge Distillation

Model ReleasesDGX agent

arXiv:2606.07474v1 Announce Type: new Abstract: Unsupervised Continual Learning (UCL) aims to enable neural networks to learn sequential tasks without labels or access to past data. A major challenge

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

Model ReleasesDGX agent

arXiv:2606.07167v1 Announce Type: cross Abstract: Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lack

VeriDrive: Verifiable Counterfactual Supervision for Cost-Efficient Vision-Language Planning

Model ReleasesDGX agent

arXiv:2606.07338v1 Announce Type: new Abstract: Vision-language driving models increasingly use reasoning supervision to bridge perception, prediction, and planning, but existing driving rationales ar

VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

Model ReleasesDGX agent

arXiv:2606.06819v1 Announce Type: new Abstract: Reasoning Video Object Segmentation (RVOS) demands a sophisticated integration of temporal dynamics, spatial details, and linguistic reasoning to achiev

we just shipped support for rubrics in deepagents ✅ give your agent a clear definition of what 'done' looks like, and force it to run in a l…

Model ReleasesDGX agent

we just shipped support for rubrics in deepagents ✅ give your agent a clear definition of what 'done' looks like, and force it to run in a loop until said goal is complete this is similar to /goal in

What Your Posts Reveal: A Benchmark and Agentic Framework for User-Level Privacy Leakage on Social Media

Model ReleasesDGX agent

arXiv:2606.06784v1 Announce Type: cross Abstract: Public social media posts can reveal private information through weak cues scattered across text, images, or metadata. Such leakage is often cumulativ

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations

Model ReleasesDGX agent

arXiv:2606.07237v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summariz

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing

Model ReleasesDGX agent

arXiv:2606.07171v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) enable flexible instruction-driven image editing, but privacy risks arise when user images expose diverse and u

When to Think Deeply: Inhibitory Deliberation for LLM Reasoning

Model ReleasesDGX agent

arXiv:2606.06745v1 Announce Type: new Abstract: Reasoning Large Language Models can improve problem-solving performance through deliberative inference, but invoking slow reasoning for every input is c

When we first demoed Claude Code internally, it got two reactions on Slack. A year after GA, @_catwu and I sat down to talk about what's cha…

Model ReleasesDGX agent

When we first demoed Claude Code internally, it got two reactions on Slack. A year after GA, @_catwu and I sat down to talk about what's changed: why I use auto mode instead of plan mode, how routines

Which Anatomy Matters Under Limited Labels? A Data-Efficient Anatomy-Aware Benchmark for Cardiac Pathology Prediction

Model ReleasesDGX agent

arXiv:2606.06509v1 Announce Type: cross Abstract: Numerous medical imaging problems must be solved under limited labels and constrained compute, yet it remains unclear whether performance gains are dr

Wordle 1,814 4/6 ⬛🟨⬛⬛⬛ ⬛⬛⬛⬛⬛ ⬛🟨🟨⬛⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result where the player solved puzzle #1,814 in four attempts, using color-coded emoji tiles (⬛ for incorrect letters, 🟨 for correct letters in wrong positions, 🟩 for

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

Model ReleasesDGX agent

arXiv:2606.06538v1 Announce Type: new Abstract: In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types

Xiaomi claims MiMo-V2.5-Pro-UltraSpeed tops 1K tokens/second, a first at the 1T-parameter scale, using a standard 8-GPU commodity node; API trial starts June 9 (Jose Antonio Lanz/Decrypt)

Model ReleasesDGX agent

Jose Antonio Lanz / Decrypt: Xiaomi claims MiMo-V2.5-Pro-UltraSpeed tops 1K tokens/second, a first at the 1T-parameter scale, using a standard 8-GPU commodity node; API trial starts June 9 — Most peop

Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs

Model ReleasesDGX agent

arXiv:2601.12359v1 Announce Type: cross Abstract: Prompt injection attacks have become an increasing vulnerability for LLM applications, where adversarial prompts exploit indirect input channels such

7 Jun 2026

A point on the board for 2026. Thank you to the team, our partners, and our fans for the ongoing work and continued support. #MonacoGP

Model ReleasesDGX agent

Aston Martin F1 team secured a point at the 2026 Monaco Grand Prix, acknowledging contributions from their team members, partners, and supporters. The post expresses gratitude for the collaborative ef

datasette-agent-edit 0.1a0

Model ReleasesDGX agent

Release: datasette-agent-edit 0.1a0 I'm planning several plugins for Datasette Agent which can make edits to existing pieces of text - things like collaborative Markdown editing, updating large SQL qu

Falcon 9 launches 21 @Starlink satellites and two Starshield satellites from California

Model ReleasesDGX agent

SpaceX's Falcon 9 rocket successfully launched 21 Starlink internet satellites and two Starshield military satellites from a California launch facility. Starlink satellites are part of SpaceX's global

Grep timeout issue fixed in latest Grok Build

Model ReleasesDGX agent

Grep timeout issue fixed in latest Grok Build Grok Build update just released v0.2.31 Release Notes: Bug Fixes: • Marketplace skills without proper descriptions are now hidden from listings instead of

Have been extensively testing Claude Workflows this weekend, with the best model possible. Threw it at my whole code base, combing for bugs.…

Model ReleasesDGX agent

Have been extensively testing Claude Workflows this weekend, with the best model possible. Threw it at my whole code base, combing for bugs. 144 found and fixed! Geez... It is a large code base, for s

LOL. Even Claude sees through Hinton’s nonsense. (though see the articles I posted earlier, for converging sources I put more weight on)

Model ReleasesDGX agent

LOL. Even Claude sees through Hinton’s nonsense. (though see the articles I posted earlier, for converging sources I put more weight on) This is my prompt and this is Claude's response: ME: What do yo

see eg https://www.scmp.com/tech/tech-trends/article/3271858/ai-race-alibaba-tencent-quickly-adopt-metas-new-llama-31-model-amid-excitement …

Model ReleasesDGX agent

see eg https://www.scmp.com/tech/tech-trends/article/3271858/ai-race-alibaba-tencent-quickly-adopt-metas-new-llama-31-model-amid-excitement and https://medium.com/the-endless-forge/zuckerbergs-llama-f

The first wave of AI-native applications is wrapping tokens and providing in-app agents. As agent usage centralizes around core apps (e.g. C…

Model ReleasesDGX agent

The first wave of AI-native applications is wrapping tokens and providing in-app agents. As agent usage centralizes around core apps (e.g. Claude Code, Codex), there's this emerging wave of building s

There was an inflection point recently where the tide shifted to model pickers and OSS Mix of tokenmaxxing/cost fatigue, nemotron coalition,…

Model ReleasesDGX agent

There was an inflection point recently where the tide shifted to model pickers and OSS Mix of tokenmaxxing/cost fatigue, nemotron coalition, harness step functions, brains/claws, etc Long live those w

Wasn't Krea 2 supposed to be released ?

Model ReleasesDGX agent

Krea 2, Krea's first foundation image model built from scratch, was announced on May 12, 2026 , with a focus on aesthetics, style transfer, and creative control . Krea 2 became available to everyone s

Wordle 1,813 4/6 ⬛⬛⬛🟨⬛ ⬛⬛⬛⬛🟨 🟨🟨🟨⬛⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result shared by Anthropic on X (Twitter), showing puzzle #1,813 solved in 4 attempts with a final correct answer displayed through colored emoji tiles (green indicat

Zuckerberg and LeCun’s unilateral decision to open source Llama likely (partly) catalyzed China’s AI industry — and may have done truly mass…

Model ReleasesDGX agent

Zuckerberg and LeCun’s unilateral decision to open source Llama likely (partly) catalyzed China’s AI industry — and may have done truly massive harm to American business interests. We are now starting

6 Jun 2026

ADK Arena: Evaluating Agent Development Kits via LLM-as-a-Developer

Model ReleasesDGX agent

arXiv:2606.05548v1 Announce Type: cross Abstract: The rapid proliferation of Agent Development Kits (ADKs), SDK-level frameworks for building LLM-powered autonomous agents, has outpaced any empirical

Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads

Model ReleasesDGX agent

arXiv:2606.06448v1 Announce Type: new Abstract: LLM agents are increasingly deployed on long-horizon tasks requiring sustained reasoning over extended interaction histories. Realizing this at scale re

Agent-Orchestrated Adaptive RAG: A Comparative Study on Structured and Multi-Hop Retrieval

Model ReleasesDGX agent

arXiv:2606.05658v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding their responses in external knowledge, but conventional pipeli

Agentic Monte Carlo: Simulating Reinforcement Learning for Black-Box Agents

Model ReleasesDGX agent

arXiv:2606.05296v1 Announce Type: cross Abstract: LLM agents operate in two distinct regimes: open-weight agents amenable to reinforcement learning (RL) and black-box agents whose behaviour must be co

AttackPathGNN: Cross-function vulnerability detection in smart contracts using state interference graphs and conjunction pooling

Model ReleasesDGX agent

arXiv:2606.05986v1 Announce Type: cross Abstract: Existing learning-based detectors for Solidity smart-contracts reduce vulnerability detection to syntactic pattern matching within single functions, y

Benchmark Everything Everywhere All at Once

Model ReleasesDGX agent

arXiv:2606.06462v1 Announce Type: new Abstract: Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their co

Benchmarking Counterfactual Prediction in Epidemic Time Series with Time-Varying Interventions

Model ReleasesDGX agent

arXiv:2606.05692v1 Announce Type: cross Abstract: Deep learning has enabled significant advances in time-series causal inference, yet progress remains constrained by the lack of realistic benchmarks w

Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillatio

Model ReleasesDGX agent

arXiv:2606.05682v1 Announce Type: new Abstract: Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost c

🚨BREAKING: Miguel Bosé, the biggest Spanish-language pop star of the last few decades, has just released a video taking a knee and putting …

Model ReleasesDGX agent

🚨BREAKING: Miguel Bosé, the biggest Spanish-language pop star of the last few decades, has just released a video taking a knee and putting his hand over his heart in honour of Henry Nowak This has now

Brick-Composer: Using MLLMs for Assembly with Diverse Bricks

Model ReleasesDGX agent

arXiv:2606.05445v1 Announce Type: new Abstract: We dream of AI agents that can read arbitrary designs and construct real-world objects from reusable building blocks. As a first step toward this vision

Can AI Refute Economic Theory? Evidence from Beyond the Knowledge Cutoff

Model ReleasesDGX agent

arXiv:2606.05383v1 Announce Type: cross Abstract: Can artificial intelligence (AI) refute economic theory? I document experiments in which I asked several AI models (Gemini, Refine, Claude, and ChatGP

Can LLMs Write Correct TLA+ Specifications? Evaluating Natural-Language-to-TLA+ Generation

Model ReleasesDGX agent

arXiv:2606.05792v1 Announce Type: new Abstract: TLA+ has supported industrial verification at companies such as Amazon and Microsoft, yet writing correct TLA+ specifications from natural language stil

CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications

Model ReleasesDGX agent

arXiv:2512.15231v3 Announce Type: replace Abstract: The automated and intelligent processing of massive remote sensing (RS) datasets is critical in Earth observation (EO). Existing automated systems a

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs

Model ReleasesDGX agent

arXiv:2606.05966v1 Announce Type: cross Abstract: Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) sti

Chip Export rules: made China/Huawei stronger, accomplished relatively little. USG stakes in US AI: will make Mistral and other “sovereign A…

Model ReleasesDGX agent

Chip Export rules: made China/Huawei stronger, accomplished relatively little. USG stakes in US AI: will make Mistral and other “sovereign AI” efforts stronger, freak out the rest of the world, escala

Closing the Loop on Latent Reasoning via Test-Time Reconstruction

Model ReleasesDGX agent

arXiv:2606.06252v1 Announce Type: new Abstract: Recent work moves intermediate reasoning from natural-language traces into latent or cache-level representations to reduce token overhead and avoid a di

CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model

Model ReleasesDGX agent

arXiv:2606.06099v1 Announce Type: new Abstract: Whether Large Language Models (LLMs) exhibit covert psychological manipulation in complex human-AI interactions has garnered increasing safety concerns.

Critic-Guided Heterogeneous Multi-Agent Reasoning for Reliable Mathematical Problem Solving

Model ReleasesDGX agent

arXiv:2606.05704v1 Announce Type: new Abstract: Recent Large Language Models (LLMs) have shown impressive reasoning abilities; but they are still susceptible to hallucinations, intermediate reasoning

Cross-Epoch Adaptive Rollout Optimization for RL Post-Training

Model ReleasesDGX agent

arXiv:2606.05606v1 Announce Type: cross Abstract: LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed ro

CTIConnect: A Benchmark for Retrieval-Augmented LLMs over Heterogeneous Cyber Threat Intelligence

Model ReleasesDGX agent

arXiv:2510.11974v2 Announce Type: replace-cross Abstract: Cyber Threat Intelligence (CTI) is foundational to modern cybersecurity, enabling organizations to proactively defend against evolving threats

Data Flow Control: Data Safety Policies for AI Agents

Model ReleasesDGX agent

arXiv:2606.05679v1 Announce Type: cross Abstract: Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users. While recent work improves query correctness,

Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows

Model ReleasesDGX agent

arXiv:2606.05670v1 Announce Type: new Abstract: Does adding more agents help an LLM workflow once compared systems share the same benchmark loader, tool access, answer contract, usage accounting, and

DPBench: Structural Determinants of Multi-Agent LLM Coordination Under Simultaneous Resource Contention

Model ReleasesDGX agent

arXiv:2602.13255v2 Announce Type: replace Abstract: We present DPBench, a benchmark for evaluating coordination in multi-agent systems built from large language models. Existing benchmarks measure tas

← Previous
1…163164165166167…377
Next →