AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,724 results
21 Jul 2026

Domesticated, not feral: Why evolvable AI is not yet a Darwinian threat Writing in PNAS (Proceedings of the National Academy of Sciences), t…

SafetyDGX agent

Domesticated, not feral: Why evolvable AI is not yet a Darwinian threat Writing in PNAS (Proceedings of the National Academy of Sciences), three heavyweights in AI and evolutionary biology criticized

Generosity Under Conditions: Hardening Google Cloud Access Management

Model ReleasesDGX agent

In Google Cloud, Identity and Access Management (IAM) helps you maintain access control over your cloud resources and operations. While it includes other features, this is its primary purpose. If you

Most startups celebrate their first couple million of revenue. @FactoryAI gave it back. They didn’t have to. They chose to. They decided tha…

Applications
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

Most startups celebrate their first couple million of revenue. @FactoryAI gave it back. They didn’t have to. They chose to. They decided that the product just wasn’t good enough yet, and they wanted t

My OCR model mislabels section titles as body text. Is a CRF the right fix, or am I overcomplicating it? [P]

Model ReleasesDGX agent

Hi everyone, I'm working on extracting the hierarchical structure of long PDF documents (legal/regulatory text, lots of numbered sections) and would like to gather some feedback on my approach before

Sakana AI is doubling down on a thesis we believe will play an increasingly important role in AI development: the future of AI won't be defi…

Model ReleasesDGX agent

Sakana AI is doubling down on a thesis we believe will play an increasingly important role in AI development: the future of AI won't be defined by a single frontier model, but by how intelligently man

20 Jul 2026

As always, Cosmos 3 Edge is fully open. This includes model weights, post-training recipes and code. Available now on @huggingface 🤗 https:…

IndustryDGX agent

NVIDIA AI has released Cosmos 3 Edge, a 4‑billion‑parameter open‑source world model designed for on‑device inference. It enables robots to learn and act, autonomous vehicles to interpret road scenes a

// Global Workspace in LLMs // arXiv paper for the popular J-space work from Anthropic. (bookmark it) The short recap: If you build on chain…

TutorialsDGX agent

// Global Workspace in LLMs // arXiv paper for the popular J-space work from Anthropic. (bookmark it) The short recap: If you build on chain-of-thought or steering vectors, this work provides a mechan

Open ecosystem or walled garden? For @anthropicai's @katelyn_lesse and @angjiang the answer is clear. 'We actually aren't precious about “Yo…

Model ReleasesDGX agent

Open ecosystem or walled garden? For @anthropicai's @katelyn_lesse and @angjiang the answer is clear. 'We actually aren't precious about “You should run these things on our infrastructure.' In practic

People are wasting their AI subscriptions. They don’t realize how powerful AI can be in achieving their own goals. They’re asking AI to repl…

Model ReleasesDGX agent

People are wasting their AI subscriptions. They don’t realize how powerful AI can be in achieving their own goals. They’re asking AI to reply to emails with zero goals in mind and nothing laddering up

16 Jul 2026

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks

Model ReleasesDGX agent

arXiv:2607.13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption across twenty

Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift

SafetyDGX agent

arXiv:2607.13221v1 Announce Type: cross Abstract: Real-time N-1 contingency screening in an energy management system trades assurance against cost: verifying every credible outage with full power flow

Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance

SafetyDGX agent

arXiv:2607.13040v1 Announce Type: cross Abstract: This paper examines where final authority should sit once capable AI systems are embedded in organizational workflows. It compares two governance mode

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

Local AiDGX agent

arXiv:2607.13960v1 Announce Type: new Abstract: World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as den

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration

Model ReleasesDGX agent

arXiv:2607.13056v1 Announce Type: cross Abstract: Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largel

LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning

SafetyDGX agent

arXiv:2607.13501v1 Announce Type: new Abstract: Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redundant, and har

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

Model ReleasesDGX agent

arXiv:2601.02023v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluctua

NVIDIAとSakana AI、オープンモデルによるイノベーションのため協業拡大 本日、Sakana AIはNVIDIAとのコラボレーションを強化し、日本発の「集合知」の取り組みを次なるフェーズへ進めることを発表します。 私たちのマルチエージェント・オーケストレーションシステム…

Model ReleasesDGX agent

NVIDIAとSakana AI、オープンモデルによるイノベーションのため協業拡大 本日、Sakana AIはNVIDIAとのコラボレーションを強化し、日本発の「集合知」の取り組みを次なるフェーズへ進めることを発表します。 私たちのマルチエージェント・オーケストレーションシステム「Sakana Fugu」に、Nemotronファミリーを含むNVIDIAのオープンモデル群を統合します。 Sakana

Privacy Preserving Recommender Systems Balancing Personalization with Privacy

SafetyDGX agent

arXiv:2607.13328v1 Announce Type: cross Abstract: Personalized recommendation systems are central to modern e-commerce and retail platforms, but they typically rely on centralized storage of detailed

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

Model ReleasesDGX agent

arXiv:2607.13039v1 Announce Type: cross Abstract: Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success. These metrics miss a d

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators

ResearchDGX agent

arXiv:2607.13075v1 Announce Type: cross Abstract: Context can change whether a request is harmful without changing its topic or surface form. We ask whether residual-stream probes distinguish harmful

We’re excited to collaborate with NVIDIA to build the next generation of Fugu orchestration models together, by incorporating leading open-w…

Model ReleasesDGX agent

We’re excited to collaborate with NVIDIA to build the next generation of Fugu orchestration models together, by incorporating leading open-weights models. Sakana AI Teams With NVIDIA to Advance Open M

15 Jul 2026

🥉 3rd place: CashFromChaos, by David Diaz (@davddiazm) CashFromChaos starts from a single seller input and automates everything up until a …

Model ReleasesDGX agent

🥉 3rd place: CashFromChaos, by David Diaz (@davddiazm) CashFromChaos starts from a single seller input and automates everything up until a completed sale. You send a photo and a one-line clue, and Her

An Empirical Study for Android-to-OpenHarmony GUI Test Migration

Model ReleasesDGX agent

arXiv:2607.11245v2 Announce Type: replace-cross Abstract: To reduce the substantial engineering effort required to test the corresponding applications from Android to OpenHarmony, migrating existing G

BREAKING: Grok 4.5 has climbed to #2 on the FrontierSWE benchmark. The result places Grok 4.5 among the world's top-performing AI models for…

Model ReleasesDGX agent

BREAKING: Grok 4.5 has climbed to #2 on the FrontierSWE benchmark. The result places Grok 4.5 among the world's top-performing AI models for software engineering tasks, highlighting its growing streng

CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.12786v1 Announce Type: new Abstract: Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attrib

Current efficient frontier of open models

Model ReleasesDGX agent

Efficiency defined as score over active parameters. Removed all the models that were not on the pareto frontier. Yes I'm aware that artificialanalysis.ai aggregate benchmark isn't perfect, but I have

DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection

Model ReleasesDGX agent

arXiv:2607.12419v1 Announce Type: new Abstract: In autonomous driving perception, the fusion of LiDAR and camera modalities has become the dominant paradigm for 3D object detection. However, current m

Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

SafetyDGX agent

arXiv:2607.12784v1 Announce Type: cross Abstract: Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However

Dynamic Resource Allocation for Ensemble Determinization MCTS

Model ReleasesDGX agent

arXiv:2607.13007v1 Announce Type: new Abstract: Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games with significant elements of randomn

Egocentric Bias in Vision-Language Models

Model ReleasesDGX agent

arXiv:2602.15892v2 Announce Type: replace-cross Abstract: Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet

Environment Parameter Gradient Theorem for Policy-Environment Co-Design in Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.12590v1 Announce Type: cross Abstract: Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, t

Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

Model ReleasesDGX agent

arXiv:2607.12739v1 Announce Type: new Abstract: A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversation

EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading

ResearchDGX agent

arXiv:2607.12455v1 Announce Type: new Abstract: Quantitative strategy optimization remains largely manual, requiring domain experts to identify weak signals, tune risk-control rules, and repeatedly va

ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning

SafetyDGX agent

arXiv:2607.12931v1 Announce Type: new Abstract: Reinforcement Learning (RL) has demonstrated significant potential for improving Vision-Language-Action (VLA) models on complex manipulation tasks. Howe

FairCoder: Probing LLM Bias in High-Stakes Decision Making via Coding Tasks

Model ReleasesDGX agent

arXiv:2501.05396v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in high-stakes decisions such as hiring and college admissions, making their social bias a critic

Good reason to try Grok 4.5 with Grok Build. It gets better every day!

Model ReleasesDGX agent

Good reason to try Grok 4.5 with Grok Build. It gets better every day! Grok 4.5 just took the #1 spot on the Long-Horizon Terminal-Bench, outperforming Claude Fable 5, Claude Opus 4.8 and GPT-5.6-sol

I’ve found @angjiang and @katelyn_lesse of @AnthropicAI to be generous, transparent, and thoughtful when it comes to building an ecosystem, …

Model ReleasesDGX agent

I’ve found @angjiang and @katelyn_lesse of @AnthropicAI to be generous, transparent, and thoughtful when it comes to building an ecosystem, not a walled garden. Listen and decide for yourself. Just so

Learning When to Trust in Contextual Social Bandits

SafetyDGX agent

arXiv:2603.13356v2 Announce Type: replace Abstract: Robust reinforcement learning typically assumes that feedback sources are either globally trustworthy or corrupted within a fixed global budget. We

LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos

ResearchDGX agent

arXiv:2607.12733v1 Announce Type: new Abstract: Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses

Mistake gating leads to energy and memory efficient continual learning

SafetyDGX agent

arXiv:2604.14336v2 Announce Type: replace Abstract: Synaptic plasticity is metabolically expensive, yet animals continuously update their internal models without exhausting energy reserves. However, w

Mobility-Aware Cache Framework for Scalable LLM-Based Human Mobility Simulation

ApplicationsDGX agent

arXiv:2602.16727v2 Announce Type: replace Abstract: Simulating large-scale human mobility is fundamental to understanding population movement patterns and supporting real-world geospatial applications

More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

Local AiDGX agent

arXiv:2607.12429v1 Announce Type: new Abstract: Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same sc

NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing

Model ReleasesDGX agent

arXiv:2603.16307v2 Announce Type: replace Abstract: Remote sensing underpins crucial applications such as disaster relief and ecological field surveys, where systems must understand complex scenes and

RL post-training on 14 Macs across 4 countries

Local AiDGX agent

Disclosure: I work at Pluralis Research, the lab that built this. Code is open, and I'm happy to answer questions. TL;DR: As far as we can tell, this is the first RL post-training run whose entire rol

Semantic-Edge Response Decoding of SAM3 for Zero-Shot Crack Segmentation

ResearchDGX agent

arXiv:2607.12292v1 Announce Type: new Abstract: Crack segmentation is essential for infrastructure inspection and structural health assessment, but existing high-performance methods typically require

🆕This Year In Claude https://www.youtube.com/watch?v=uU5Gv2h8-9g @simonw chats with @_catwu and @trq212 about the state of: - @claudeai Cod…

Model ReleasesDGX agent

🆕This Year In Claude https://www.youtube.com/watch?v=uU5Gv2h8-9g @simonw chats with @_catwu and @trq212 about the state of: - @claudeai Code - Claude Fable - @anthropicai culture & product strategy -

Training against GPT‑Red makes GPT‑5.6 substantially more resilient. To measure this, we replayed some of GPT‑Red’s strongest attacks—none o…

ApplicationsDGX agent

Training against GPT‑Red makes GPT‑5.6 substantially more resilient. To measure this, we replayed some of GPT‑Red’s strongest attacks—none of which our models had seen during training. GPT‑5.6 Sol pro

Transforming LLMs into Efficient Cross-Encoders via Knowledge Distillation for RAG Reranking

Model ReleasesDGX agent

arXiv:2607.11933v1 Announce Type: new Abstract: Cross-encoders achieve high reranking accuracy in Retrieval-Augmented Generation (RAG) pipelines but impose quadratic inference costs that limit real-ti

Vertical Standardisation for High-Risk AI Systems under the EU AI Act: A Domain-Specific Framework for Algorithmic Hiring

SafetyDGX agent

arXiv:2607.12588v1 Announce Type: new Abstract: According to the recent European legislation, high-risk AI systems will have to adapt in order to comply with requirements related to specific areas, li

WanToFight: Real-Time Generative Game Engine for Multi-Player Combat Interaction

Local AiDGX agent

arXiv:2607.12592v1 Announce Type: new Abstract: We present WanToFight, a generative game engine that simulates real-time, two-player The King of Fighters '97 (KOF~'97) gameplay from keyboard input. Pr

14 Jul 2026

Cribl buys AI-native threat detection engineering startup CardinalOps

ApplicationsDGX agent

Data observability and telemetry management company Cribl Inc. says it’s extending its capabilities into security operations after acquiring the artificial intelligence-native security engineering sta

Fortinet expands FortiEndpoint with AI monitoring and data controls

IndustryDGX agent

Fortinet Inc. today expanded its FortiEndpoint platform with a set of features aimed at helping companies monitor employee use of artificial intelligence tools and stop sensitive data from leaking int

Kimi K3 in the next few hours. Deepseek V4 GA later in the week. New Liquid models. New Mistral models sometime this month. And some rumours suggest GLM 5.5 is coming in August. Openweight AI is eating good.

Model ReleasesDGX agent

dam bois we eating good this week ngl, The velocity of the open_weight ecosystem right now is hitting a point where proprietary, closed-source APIs are losing their leverage on compute intelligence. W

Today, we’re announcing Bonsai 27B: the first 27B-class model to run on a phone. Bonsai 27B is the new multimodal flagship of the Bonsai fam…

Local AiDGX agent

Today, we’re announcing Bonsai 27B: the first 27B-class model to run on a phone. Bonsai 27B is the new multimodal flagship of the Bonsai family. Based on Qwen3.6 27B, it brings a new capability tier t

We’re open sourcing WANDR. WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexi…

Model ReleasesDGX agent

We’re open sourcing WANDR. WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer. https://research.perplexity.ai/articles/wandr-b

What are the best models you can run on your @NVIDIAAI DGX Spark? ✨ Mid-July 2026 Edition 1× DGX Spark • ⁠Qwen 3.6 35b NVFP4 — 256k ctx, 81 …

Model ReleasesDGX agent

What are the best models you can run on your @NVIDIAAI DGX Spark? ✨ Mid-July 2026 Edition 1× DGX Spark • ⁠Qwen 3.6 35b NVFP4 — 256k ctx, 81 tok/s • ⁠Qwen 3.6 27b NVFP4 — 256k ctx, 33 tok/s 2× DGX Spar

13 Jul 2026

Standard RL benchmarks are episodic and stationary, so they don't capture the the characteristics of real-world deployment. Morpheus is a ne…

Model ReleasesDGX agent

Standard RL benchmarks are episodic and stationary, so they don't capture the the characteristics of real-world deployment. Morpheus is a new benchmark for continual learning that provides persistent

12 Jul 2026

I don't think Anthropic realizes how disruptive these changes are to users. I appreciate the extension, but please stop playing games. Eithe…

Model ReleasesDGX agent

I don't think Anthropic realizes how disruptive these changes are to users. I appreciate the extension, but please stop playing games. Either keep it under the subscriptions or put it under the API al

11 Jul 2026

VLMは人間のような創造性を持てるか? ケネス・スタンレー教授らの『目標という幻想(Why Greatness Cannot Be Planned)』は、明確な目標を設定することが、かえって真に偉大な発見を遠ざけてしまうという逆説を論じた書籍です。その議論の中核にあったのが「Pi…

ResearchDGX agent

VLMは人間のような創造性を持てるか? ケネス・スタンレー教授らの『目標という幻想(Why Greatness Cannot Be Planned)』は、明確な目標を設定することが、かえって真に偉大な発見を遠ざけてしまうという逆説を論じた書籍です。その議論の中核にあったのが「PicBreeder」の実験でした。 PicBreeder では、ユーザーが「面白い」と感じた画像を選び、それを少しずつ進化

10 Jul 2026

ADORN: Adaptive Drift handling for Open RAN using Reinforcement Learning

SafetyDGX agent

arXiv:2607.08443v1 Announce Type: cross Abstract: Dynamic traffic variations in Open Radio Access Networks (O-RAN) lead to drift, which degrades the performance of Artificial Intelligence/Machine Lear

← Previous
1…267268269270271…296
Next →