AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
13 Aug 2026

Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

Model ReleasesDGX agent

arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers

Grok 4.6 ranks #1 on the GPQA Diamond leaderboard 🧠 Grok 4.6 (high) scores 95% - the highest score on the chart for graduate-level scientif…

Model ReleasesDGX agent

Grok 4.6 ranks #1 on the GPQA Diamond leaderboard 🧠 Grok 4.6 (high) scores 95% - the highest score on the chart for graduate-level scientific reasoning It outperforms Claude Fable 5, Opus 5, GPT-5.6 S

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.12063v1 Announce Type: cross Abstract: Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely

OEIS Open: How many conjectures can language models turn into theorems?

Model ReleasesDGX agent

arXiv:2608.11941v1 Announce Type: new Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conje

out of curiosity, i tasked fable to figure this thing out and it looks like the 'saas killer' is done. ▲ it maxed out Fable 5 + used most of…

Model ReleasesDGX agent

out of curiosity, i tasked fable to figure this thing out and it looks like the 'saas killer' is done. ▲ it maxed out Fable 5 + used most of Codex limits ▲ most of the work was done within 45h session

To use Computer History, opt in under Settings → Integrations in the ChatGPT desktop app on Mac. Rolling out globally now to Pro, Business, …

ApplicationsDGX agent

To use Computer History, opt in under Settings → Integrations in the ChatGPT desktop app on Mac. Rolling out globally now to Pro, Business, and Enterprise users, with access in the EEA, UK, and Switze

Try Grok 4.6

Model ReleasesDGX agent

Try Grok 4.6 Grok 4.6 wins again. 👑 Grok 4.6 takes the #1 spot on GPQA Diamond with a score of 94.9%, beating GPT-5.6, Gemini 3.1 Pro, Claude Opus 5, and every other model tested by Artificial Analysi

Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits

SafetyDGX agent

arXiv:2608.11410v1 Announce Type: new Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Err

12 Aug 2026

A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language

Model ReleasesDGX agent

arXiv:2608.10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically

An adaptive and evolvable deep reinforcement learning framework for weather prediction

Model ReleasesDGX agent

arXiv:2608.09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times. Rather than building yet another architecture, we reframe the for

Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

Model ReleasesDGX agent

arXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation

Model ReleasesDGX agent

arXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

SafetyDGX agent

arXiv:2608.10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world mod

ELMER: Evolutionary Language Model that Explores and Refines

SafetyDGX agent

arXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is a

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering

Model ReleasesDGX agent

arXiv:2608.10679v1 Announce Type: cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are wor

FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition

Model ReleasesDGX agent

arXiv:2608.10396v1 Announce Type: new Abstract: Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure

How to Choose Full-Stack Observability for NVIDIA AI Factories

HardwareDGX agent

A full‑stack observability framework for NVIDIA AI factories links telemetry from compute, networking, storage, orchestration and application layers using specialized tools (DCGM, NVSM, UFM, NetQ, NMX

IO Factory: Simulating AI-Enabled Influence Campaigns at Scale

SafetyDGX agent

arXiv:2608.10920v1 Announce Type: new Abstract: We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat

MindTopo reveals VLMs’ spatial reasoning abilities

Model ReleasesDGX agent

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post

Quoting Florian Herrengt

Model ReleasesDGX agent

But then users start to report a weird bug. It's the 4th time your team has been trying to fix it. I mean... asking AI to fix it. Unfortunately, it seems like not even Fable can figure it out. You go

Qwen 3.8 2.4T is out , no 27b today RIP.

Model ReleasesDGX agent

i was hoping they would release both today but the big one just dropped now , and it says its been listed since 5 hours ago on huggingface. I guess the real date for 27b is https://modelscope.cn/model

Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator

Model ReleasesDGX agent

arXiv:2608.09971v1 Announce Type: cross Abstract: Over the past few years, the rapid development of machine learning (ML) models for weather forecasting has produced deterministic models whose medium-

Sources detail moves behind Google's AI reshuffle; Sergey Brin urged key staff to go all in on Gemini, and some teams shifted from DeepMind to corporate Google (Kenrick Cai/Reuters)

Model ReleasesDGX agent

Kenrick Cai / Reuters: Sources detail moves behind Google's AI reshuffle; Sergey Brin urged key staff to go all in on Gemini, and some teams shifted from DeepMind to corporate Google — Google co-found

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

Model ReleasesDGX agent

arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached o

V-FiLLM: Verified Financial LLM Reasoning Benchmark

Model ReleasesDGX agent

arXiv:2608.11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains compar

We wrote a 36-page ArXiv whitepaper on ExtractBench 🧑‍🔬 , our effort to create the most comprehensive, schema-guided, real-world document …

Model ReleasesDGX agent

We wrote a 36-page ArXiv whitepaper on ExtractBench 🧑‍🔬 , our effort to create the most comprehensive, schema-guided, real-world document extraction benchmark. It’s extremely detailed and covers every

Why does CLAUDE.MD keep growing? If you maintain a CLAUDE.md or an AGENTS.md, this one is worth your time. (bookmark it) This work traces wh…

Model ReleasesDGX agent

Why does CLAUDE.MD keep growing? If you maintain a CLAUDE.md or an AGENTS.md, this one is worth your time. (bookmark it) This work traces why these files grow without bound. Appending an instruction i

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

SafetyDGX agent

arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream dec

11 Aug 2026

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

Model ReleasesDGX agent

arXiv:2607.10180v2 Announce Type: replace-cross Abstract: We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. T

AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

Model ReleasesDGX agent

arXiv:2608.07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist en

Anyone Using (Koreas) 'Solar Open 2' (250B, 15B) Model?

Model ReleasesDGX agent

I just heard of this model. Seems to be a competitor to DeepSeek V4 Flash. About the same size and active parameters. Anyone tested it compared to V4 Flash? Link: https://huggingface.co/upstage/Solar-

Back to the Future: A workbook time machine for spread sheet creation benchmarks

Model ReleasesDGX agent

arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived obj

Blumira launches Hearth, an AI command center that spans rival security tools

Model ReleasesDGX agent

Security operations platform startup Blumira Inc. today launched Hearth, a vendor-agnostic artificial intelligence command center that lets security teams investigate and act across their existing too

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2608.08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the 'right' answer, but how they decide what matters most when moral principle

DeepSeek-V4-Flash-0731 (284B MoE) at 75 tok/s on 2× DGX Spark — full recipe, 11 gotchas, reboot-proof cluster, Codex CLI integration

Model ReleasesDGX agent

Spent two nights getting deepseek-ai/DeepSeek-V4-Flash-0731 (284B MoE, 13B active, native FP4/FP8, 1M context) running production-grade on two DGX Sparks connected by one QSFP DAC cable. Everything —

EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility

Model ReleasesDGX agent

arXiv:2608.08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity

HarnessWAM: Bridging Prediction and Deliberation in World Action Models

Local AiDGX agent

arXiv:2608.09516v1 Announce Type: new Abstract: World Action Models (WAMs) jointly learn environmental dynamics and robot actions, introducing priors over physical evolution into embodied control. How

I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

Model ReleasesDGX agent

I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon

Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue

SafetyDGX agent

arXiv:2608.08210v1 Announce Type: new Abstract: Collaborative dialogue can end with apparent agreement while participants still differ on goals, assumptions, or execution plans, creating an extbf{illu

Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models a…

Model ReleasesDGX agent

Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models are pushing the frontier of coding and knowledge work, but su

LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling

Model ReleasesDGX agent

arXiv:2608.09343v1 Announce Type: new Abstract: Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a black box: aggreg

ML-Based Hierarchical Prediction for Practical Energy Scheduling in Dynamic NTN-WPT Systems

SafetyDGX agent

arXiv:2608.08804v1 Announce Type: cross Abstract: With advancements in long-distance wireless power transfer (WPT) and space-based energy technologies, integrating WPT into non-terrestrial networks (N

Observations on Muse-Glimmer reasoning traces being noticeably different from qwen / gemma models and questions for you guys

Model ReleasesDGX agent

Just downloaded the model, UD-Q5_K_XL quant, asked it to generate a long story to test out reasoning and speed with dflash (super fast btw, ~ 90 to 160 tok/s on a 5090 depending on task) and was surpr

Open source is so back. Zuck just announced Meta is opening the weights for Muse Glimmer, with Muse Spark 1.2 coming soon But a year ago, ev…

Model ReleasesDGX agent

Open source is so back. Zuck just announced Meta is opening the weights for Muse Glimmer, with Muse Spark 1.2 coming soon But a year ago, everyone doubted Meta's position in the AI race In an intervie

OpenWALDO launches to build collaborative community for open-source AI

Model ReleasesDGX agent

OpenWALDO, a new open-source artificial intelligence project sponsored by Ctrl IQ Inc., launched today, led by Gregory Kutzer, the founder of Rocky Linux, CentOS and Apptainer. The project aims to bui

Personalized AI startup River AI raises $1.1B from consortium backed by Nvidia, AMD

HardwareDGX agent

River AI Inc., a startup that helps enterprises customize open-source artificial intelligence models, has raised 1.1 billion in early-stage funding. The company stated in today’s announcement that it

The Anatomy of a Prompt Injection: A Component Model for Structured Analysis

SafetyDGX agent

arXiv:2608.07808v1 Announce Type: cross Abstract: Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than structured

💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to c…

Model ReleasesDGX agent

💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to choose the right model for the right task. As part of this, we

Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days. Their Meta Muse Glimmer m…

Model ReleasesDGX agent

Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days. Their Meta Muse Glimmer model is a 30B multimodal reasoning model with a Gemma-like a

10 Aug 2026

An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2608.07439v1 Announce Type: new Abstract: Quantum natural language processing (QNLP) provides a grammar-aware framework for text modeling, and Distributional Compositional Categorical (DisCoCat)

Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering

Model ReleasesDGX agent

arXiv:2608.07038v1 Announce Type: cross Abstract: Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereby reducing

Chat UIs with native audio input for multimodal models?

Model ReleasesDGX agent

I've been running Gemma 4 E4B with oMLX and I can't find any chat interfaces that directly send the audio file to the model instead of running the audio through a separate STT layer. I can confirm the

CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing

Model ReleasesDGX agent

arXiv:2608.07424v1 Announce Type: new Abstract: Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or applying a str

Corma launches with $60M in funding for defensive cybersecurity AI

Model ReleasesDGX agent

Defensive cybersecurity startup Corma Labs Ltd. today announced it has raised 60 million in seed funding to build a foundation model purpose-built for security defense. Founded in 2025, Corma runs off

DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding

ResearchDGX agent

arXiv:2608.07067v1 Announce Type: new Abstract: Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static

FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

Model ReleasesDGX agent

arXiv:2608.07400v1 Announce Type: new Abstract: Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be gro

GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

Model ReleasesDGX agent

arXiv:2608.07411v1 Announce Type: new Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into

Got Meta's new Muse Glimmer 30B running on my MacBook (M3 Max, 96GG) and tested the serving options available so far. Fastest right now: Oll…

Model ReleasesDGX agent

Got Meta's new Muse Glimmer 30B running on my MacBook (M3 Max, 96GG) and tested the serving options available so far. Fastest right now: Ollama's MLX engine (DFlash included) at ~29 tok/s. Tuned llama

How Malachyte solves retail’s cold-start problem with managed real-time AI

Local AiDGX agent

What’s the best way to recommend products to little-known users? We’ve spent our careers trying to solve this problem for major companies like Spotify and Priceline, and it’s why Sidd founded Malachyt

How WPP operationalizes platform and data engineering for AI marketing

SafetyDGX agent

Between chaotic levels of market fragmentation and economic volatility, marketing and communications agencies can no longer rely on the human intuition they’ve traditionally used to win clients and op

← Previous
1…240241242243244…297
Next →