AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
24 Jul 2026

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Model ReleasesDGX agent

arXiv:2607.21072v1 Announce Type: new Abstract: Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks a

Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI

AgentsDGX agent

arXiv:2607.20916v1 Announce Type: new Abstract: Generative AI lets large language models produce scholarly-looking text within seconds, yet fluency does not equal valid explanation. The deepest risk i

23 Jul 2026

AI infrastructure demand is outrunning even the boldest supply chain playbooks

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents
DGX agent

AI infrastructure buildouts are moving so fast that plans made just months ago are already obsolete, forcing hardware makers to rewrite how they design, source and ship the systems powering the next g

AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic

AgentsDGX agent

arXiv:2607.11338v2 Announce Type: replace Abstract: Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is challenging. T

DeepSeek Founder’s 4-hour investor meeting: DeepSeek is prioritizing AGI over user growth and commercialisation

Model ReleasesDGX agent

A Chinese article compiled 52 remarks from Liang Wenfeng’s four-hour investor meeting. I’ve summarised the most important ones below. DeepSeek has one central objective: AGI. This is not the time to m

Going to “Advancing AI” by @AMD with @LisaSu this afternoon. Good thing they didn’t forget a space in the topic of this chat 😂

AgentsDGX agent

On July 23, 2026, the user posted a tweet announcing attendance at the “Advancing AI” event hosted by AMD in San Francisco, where Lisa Su would be speaking. The tweet noted the importance of a proper

PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code Optimization

Model ReleasesDGX agent

arXiv:2607.19653v1 Announce Type: cross Abstract: Large language model (LLM) agents now perform well on correctness-oriented repository-level tasks, including SWE-Bench issue resolution and feature im

Predictive single cell foundation model for gene regulation and aging with privacy-preserving tabular learning

AgentsDGX agent

arXiv:2607.19400v1 Announce Type: new Abstract: Pre-trained foundation models (FMs) have begun transforming single-cell genomics, but scaling them raises privacy concerns. Moreover, unlike text data,

22 Jul 2026

I don’t believe reality is a simulation, but you genuinely couldn’t script this timeline: • Two weeks ago: At @swyx’s AI Engineer World’s Fa…

AgentsDGX agent

I don’t believe reality is a simulation, but you genuinely couldn’t script this timeline: • Two weeks ago: At @swyx’s AI Engineer World’s Fair in SF, I decide at the last minute to introduce my friend

You can talk to Grok like a person to accomplish tasks via Grok Build http://X.ai/cli

AgentsDGX agent

You can talk to Grok like a person to accomplish tasks via Grok Build http://X.ai/cli I’ve gotten so used to Grok’s speech-to-text that typing now feels like hell Quick tip: Even when you’re using ano

21 Jul 2026

How OpenAI uses human feedback to evaluate and improve LLMs

AgentsDGX agent

At ChatGPT scale, user frustration arrives as support tickets, ratings, social posts, and corrections buried inside conversations. OpenAI built a feedback system that can find the pattern behind a com

My 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out! 1. Closed vs open models 2. Throughput maxxing but accur…

Model ReleasesDGX agent

My 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out! 1. Closed vs open models 2. Throughput maxxing but accuracy minimizing 3. Benchmaxxing & cheating 4. Distillation &

Using Ollama as a server

Model ReleasesDGX agent

I am currently running Qwen3.6-30B in Ollama, through Cline to use as an agent in VSCode. Qwen's skill in coding is not in question, but the performance in VSCode is slow and inaccurate and times out

16 Jul 2026

Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis

Model ReleasesDGX agent

arXiv:2607.13602v1 Announce Type: cross Abstract: Systematic comparisons between current situations and structurally similar past events in the historical, i.e., historical analogies, is among the mos

Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review

AgentsDGX agent

arXiv:2607.13054v1 Announce Type: cross Abstract: Environmental monitoring with unmanned aerial vehicles (UAVs) requires route planning methods that maximize covered area while handling energy limits,

HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models

Local AiDGX agent

arXiv:2607.13072v1 Announce Type: cross Abstract: Zero-shot object-goal navigation aims to enable an intelligent agent to explore and navigate to objects of unknown categories in an unfamiliar environ

It is clear open source models and harnesses are having a moment. There's a few factors at work 1/ It is now obvious that you can catch up t…

AgentsDGX agent

It is clear open source models and harnesses are having a moment. There's a few factors at work 1/ It is now obvious that you can catch up to near-SOTA performance and do so with a clear training line

Lyapunov Exponent as Physics-Informed Dense Reward: RL Discovery of Stabilization Beyond the Kapitza Pendulum

AgentsDGX agent

arXiv:2607.14001v1 Announce Type: new Abstract: We suggest using the Lyapunov characteristic exponent (LCE) as a dense reward signal for the reinforcement learning problem of stabilizing the inverted

RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar

Model ReleasesDGX agent

arXiv:2607.13189v1 Announce Type: cross Abstract: We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chin

Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation

AgentsDGX agent

arXiv:2607.14006v1 Announce Type: cross Abstract: Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, configurations, or operational con

Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild

AgentsDGX agent

arXiv:2607.13881v1 Announce Type: cross Abstract: Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction categories.

15 Jul 2026

A model drop by Thinky 🚨🚨 Have been doing some early testing on the model for the past couple of days. Here are some of my findings 1. The…

AgentsDGX agent

A model drop by Thinky 🚨🚨 Have been doing some early testing on the model for the past couple of days. Here are some of my findings 1. The reasoning is sharp and concise! Always love to see models tha

as someone who does a lot of dev community and dev youtube this pace of growth has been one of the greatest mysteries to me because clearly …

AgentsDGX agent

as someone who does a lot of dev community and dev youtube this pace of growth has been one of the greatest mysteries to me because clearly there's something to learn here @sytses i'd love to talk to

Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)

Model ReleasesDGX agent

Below Upstream Status sections are from https://github.com/PrismML-Eng/Bonsai-demo Upstream Status for Binary Q1_0 is supported out of the box in upstream llama.cpp across many backends: CPU (generic,

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

Model ReleasesDGX agent

arXiv:2607.12252v1 Announce Type: new Abstract: Deep research agents are increasingly used to produce long-form financial reports, yet large-scale evaluation remains bottlenecked by the need for human

Inkling is our first open model from @thinkymachines and is now available on Tinker! Check out these quotes from Tinker customers on their e…

AgentsDGX agent

Inkling is our first open model from @thinkymachines and is now available on Tinker! Check out these quotes from Tinker customers on their experience with Inkling: @_Mantic_AI: 'Not only does Inkling

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement

AgentsDGX agent

arXiv:2606.19387v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved remarkable success in software development. However, they are susceptible to hallucinations, meanin

Practical Judgment, Virtue, and Intuition in the Use of Opaque AI-Enabled Systems

AgentsDGX agent

arXiv:2607.12755v1 Announce Type: cross Abstract: AI-enabled systems are seeing increasing deployment across numerous domains, with many being 'black boxes' with respect to core functions and capabili

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

Model ReleasesDGX agent

arXiv:2607.12477v1 Announce Type: new Abstract: Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenar

Something I have been thinking about: in the past, the best engineers I knew spent a lot of time automating their work in various ways. Bett…

Model ReleasesDGX agent

Something I have been thinking about: in the past, the best engineers I knew spent a lot of time automating their work in various ways. Better vim/emacs automations, writing lint rules to catch repeat

Traj-VLN: Learning Pixel-Space Interaction via Autoregressive Trajectory Generation

AgentsDGX agent

arXiv:2607.10744v2 Announce Type: replace Abstract: Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models

xai-org/grok-build, now open source

Model ReleasesDGX agent

xai-org/grok-build, now open source xAI's grok CLI tool faced severe community backlash yesterday when it became apparent that running the command in a directory could upload that entire directory to

14 Jul 2026

Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency

AgentsDGX agent

Power is AI infrastructure’s inescapable constraint. How many tokens an AI factory can generate within a fixed power budget determines its revenue and profitability. Because of this, performance per w

10 Jul 2026

ArtMine: Discovering and Formalizing Artistic Processes

AgentsDGX agent

arXiv:2607.08331v1 Announce Type: cross Abstract: Understanding how artworks are created requires reasoning about the iterative decisions, material operations, and contextual influences that shape art

great overview of general purpose brain!

AgentsDGX agent

great overview of general purpose brain! I just published a new video on OpenWiki Brains, specifically on the general-purpose brain mode. In the video I dive into: - configuring it locally - its archi

Here is a visualization of the AI Picbreeder engine in action. (From our blog: https://pub.sakana.ai/picbreeder-vlm/) To recreate a collabor…

ApplicationsDGX agent

Here is a visualization of the AI Picbreeder engine in action. (From our blog: https://pub.sakana.ai/picbreeder-vlm/) To recreate a collaborative human ecosystem, we run 10 VLM “breeder” agents in par

I just published a new video on OpenWiki Brains, specifically on the general-purpose brain mode. In the video I dive into: - configuring it …

AgentsDGX agent

I just published a new video on OpenWiki Brains, specifically on the general-purpose brain mode. In the video I dive into: - configuring it locally - its architecture - what the docs look like - new f

𝕏 is a great platform for product announcements, especially if done by the CEO directly. Way more interesting to the public than generic pr…

AgentsDGX agent

𝕏 is a great platform for product announcements, especially if done by the CEO directly. Way more interesting to the public than generic press releases. This post by Mark Zuckerberg already received o

Keeping up with AI news is becoming a full-time job. So my friend @ivan_bezdomny built HuggingNews, an AI-curated feed that surfaces the new…

AgentsDGX agent

Keeping up with AI news is becoming a full-time job. So my friend @ivan_bezdomny built HuggingNews, an AI-curated feed that surfaces the news actually worth reading. Soon, it will even personalize the

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

AgentsDGX agent

arXiv:2607.08182v1 Announce Type: cross Abstract: Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic

ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation

AgentsDGX agent

arXiv:2607.08691v1 Announce Type: cross Abstract: Repository-level code generation requires implementing target functions while accounting for complex cross-file dependencies and project-specific conv

Self-Adaptive Anomaly Detection with Reinforcement Learning and Human Feedback in Connected Vehicles

AgentsDGX agent

arXiv:2607.08373v1 Announce Type: cross Abstract: Connected vehicles are autonomous cyber-physical systems whose behavior must be continuously monitored during operation to detect deviations from norm

Sources: Manus investors and management are discussing unwinding Meta's $2B buyout at the same valuation, with Tencent in talks to become the largest investor (Zijing Wu/Financial Times)

AgentsDGX agent

Zijing Wu / Financial Times: Sources: Manus investors and management are discussing unwinding Meta's $2B buyout at the same valuation, with Tencent in talks to become the largest investor — Chinese te

The most important thing about Grok Build and the 4.5 release is that it is genuinely so useful for real-world work

AgentsDGX agent

The most important thing about Grok Build and the 4.5 release is that it is genuinely so useful for real-world work Grok 4.5 just topped Perplexity’s WANDR orchestrator evaluation It scored higher tha

When Does Continual Learning Require Learning

AgentsDGX agent

arXiv:2607.07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn? Today, the field largel

9 Jul 2026

Gimitest: A Comprehensive Tool for Testing Reinforcement Learning Policies

AgentsDGX agent

arXiv:2607.07029v1 Announce Type: cross Abstract: Reinforcement learning (RL) policies can be unsafe and vulnerable to attacks. Ensuring their reliability is often a pain point as existing automated t

Grok 4.5 on OpenClaw

AgentsDGX agent

Grok 4.5 on OpenClaw Grok 4.5 from @SpaceXAI is live on OpenClaw. No OpenClaw update required, just connect your X Premium or SuperGrok subscription, select Grok 4.5 under the xAI provider, and use an

Measuring Intelligence Beyond Human Scale

AgentsDGX agent

arXiv:2607.07040v1 Announce Type: new Abstract: How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which ta

My feed has been hijacked this week by frontier model influencers who say they've had 5.6 Sol and Fable for 'months'. I think this tells an …

Model ReleasesDGX agent

My feed has been hijacked this week by frontier model influencers who say they've had 5.6 Sol and Fable for 'months'. I think this tells an inaccurate story the field. From the outside it makes AI pro

Neutral Substrates: A Design Constraint for Shared Records Under Persistent Interpretive Disagreement

AgentsDGX agent

arXiv:2601.14271v2 Announce Type: replace Abstract: Shared accountability records are often used by parties who may never agree about causation, responsibility, or normative interpretation. For such r

Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context

AgentsDGX agent

Long-context handling remains a core challenge for language models: even with extended context windows, models often fail to reliably extract, reason over, and use the information across long contexts

Thank you @hwchase17 @BraceSproul @devstein64 @jeffreyhuber for the LLM Wikis session today, I absolutely loved it! Really grateful you keep…

AgentsDGX agent

Thank you @hwchase17 @BraceSproul @devstein64 @jeffreyhuber for the LLM Wikis session today, I absolutely loved it! Really grateful you keep these sessions open and honest about what actually works in

Token per watt becomes the defining metric as storage moves to AI’s critical path

AgentsDGX agent

Token per watt — not raw compute — is emerging as the defining efficiency metric for AI data centers, putting storage at the center of an infrastructure rethink that is reshaping how the industry meas

8 Jul 2026

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems

AgentsDGX agent

arXiv:2607.05638v1 Announce Type: cross Abstract: Teams deploying large language models in business contexts need evaluation systems, yet most treat evaluation as static model selection: run benchmark

Evaluating calibrated refusal and safe usefulness in dual-use biology settings

Model ReleasesDGX agent

arXiv:2607.05462v1 Announce Type: cross Abstract: As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refu

Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search

AgentsDGX agent

arXiv:2607.05970v1 Announce Type: cross Abstract: Dataset search depends heavily on metadata, making LLM-generated metadata a consequential form of synthetic content in retrieval systems. We study six

From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space

SafetyDGX agent

arXiv:2607.05794v1 Announce Type: new Abstract: Long-term user memory is essential for personalized conversational agents, yet many memory systems still expose memory through passive retrieval interfa

Learning The Minimum Action Distance

AgentsDGX agent

arXiv:2506.09276v4 Announce Type: replace-cross Abstract: This paper presents a state representation framework for Markov decision processes (MDPs) that can be learned solely from state trajectories,

Quantifying Frontier LLM Capabilities for Container Sandbox Escape

Model ReleasesDGX agent

arXiv:2603.02277v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks, cr

Rewriting Bun in Rust

Model ReleasesDGX agent

Rewriting Bun in Rust Jarred Sumner has been promising this blog post (since May 9th) about his Zig to Rust rewrite of Bun for significantly longer than it took him to finish the rewrite. Honestly, it

← Previous
1…183184185186187…300
Next →