AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,737 results
12 Aug 2026

imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here …

Model ReleasesDGX agent

imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here to top it at a great price, congrats to @SpaceXAI team & loo

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

SafetyDGX agent

arXiv:2608.10635v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain chall

Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.10050v1 Announce Type: new Abstract: Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business ch

Optimal Stopping of Self-Refining Foundation Models

Model ReleasesDGX agent

arXiv:2608.10729v1 Announce Type: cross Abstract: Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is embedded in a

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

ResearchDGX agent

arXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud

Predicting Space Groups of Double Perovskites by LLM with Dynamic Few-Shot Learning

SafetyDGX agent

arXiv:2608.10483v1 Announce Type: new Abstract: Double perovskites (DPs) offer broad compositional tunability, but predicting the space groups (SGs) of stable structures remains difficult because avai

Procedural Fairness Failures in RLHF from Preference Averaging

SafetyDGX agent

arXiv:2608.10126v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. Wh

Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory

Local AiDGX agent

arXiv:2608.11066v1 Announce Type: cross Abstract: We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a future-acce

Qwen3.8-Max is live on Together AI on day zero. Proud to have supported early testing alongside @Alibaba_Qwen, @vllm_project, and Inferact t…

Model ReleasesDGX agent

Qwen3.8-Max is live on Together AI on day zero. Proud to have supported early testing alongside @Alibaba_Qwen, @vllm_project, and Inferact to get day-zero support live. We’re continuing to optimize pe

R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

Local AiDGX agent

arXiv:2608.11017v1 Announce Type: cross Abstract: Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it la

ReLTEx: Reliable LLM-based Taxonomy Expansion

Model ReleasesDGX agent

arXiv:2608.10970v1 Announce Type: cross Abstract: Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, maki

Rethinking Text-Based Image Retrieval in Specific Domain

Model ReleasesDGX agent

arXiv:2608.10524v1 Announce Type: cross Abstract: Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, exis

Scheduling Mixed RL Rollouts Beyond Prefix Locality

SafetyDGX agent

arXiv:2608.11152v1 Announce Type: cross Abstract: Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple dom

Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72

Model ReleasesDGX agent

Alibaba released the open‑weights Qwen3.8‑2.4T‑A95B (Qwen3.8‑Max), a fine‑grained mixture‑of‑experts model with 2.4 trillion parameters, hybrid full‑ and linear‑attention, a one‑million‑token context

SPACEXAI: Grok 4.6 leads on the two strongest knowledge-work / real-world productivity benchmarks (GDPVal-AA and AA-Briefcase) and on the le…

Model ReleasesDGX agent

SPACEXAI: Grok 4.6 leads on the two strongest knowledge-work / real-world productivity benchmarks (GDPVal-AA and AA-Briefcase) and on the legal benchmark, while remaining highly competitive on coding-

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

ResearchDGX agent

arXiv:2608.10949v1 Announce Type: cross Abstract: Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under

The most dangerous document extraction failure isn't a wrong value. It's a missing row that looks like nothing is wrong. We released Extract…

Model ReleasesDGX agent

The most dangerous document extraction failure isn't a wrong value. It's a missing row that looks like nothing is wrong. We released ExtractBench yesterday: 370 enterprise docs, 14 systems. The hardes

Toward a Theory of Value in AI Alignment

SafetyDGX agent

arXiv:2608.10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms

Twitch streamers can now opt out from training Amazon’s AI

IndustryDGX agent

Twitch users can now opt out of allowing their content to be used to train Amazon's generative AI models. Opting out means that 'your streams, VODs, clips, stream chats, and pictures and text on your

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

Model ReleasesDGX agent

arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of

Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

Model ReleasesDGX agent

arXiv:2608.11022v1 Announce Type: cross Abstract: Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their con

11 Aug 2026

1 Day in and I feel okay saying Muse-Glimmer-30B finally beats 3.6-27B for the size in some use-cases

Model ReleasesDGX agent

A few things right off the bat: it reasons very efficiently. Like Grok 4.5 levels of efficient thinking it quantizes very well. My first few tests with iq3_xxs were better than Qwen/Gemma behaved at t

A foundation model of numerical intelligence with cross-disciplinary generalization

ResearchDGX agent

arXiv:2607.28432v2 Announce Type: replace Abstract: Intelligence is commonly understood as the ability to acquire and apply knowledge, adapt to unfamiliar situations and solve new problems. Large lang

Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks

Model ReleasesDGX agent

arXiv:2603.00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduc

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

Model ReleasesDGX agent

arXiv:2608.07584v1 Announce Type: new Abstract: Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such t

Cross-Model Humor Preference Modeling with Cards Against Humanity

Model ReleasesDGX agent

arXiv:2608.07481v1 Announce Type: cross Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style

DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving

SafetyDGX agent

arXiv:2608.09333v1 Announce Type: new Abstract: Large-scale language models for autonomous driving enable enhanced global understanding and long-horizon planning. However, when deployed in isolated ve

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

SafetyDGX agent

arXiv:2608.09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition

SafetyDGX agent

arXiv:2608.09762v1 Announce Type: new Abstract: Real-world online reinforcement learning (RL) provides a promising approach for training robotic manipulation policies directly in the physical world, a

ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making

Model ReleasesDGX agent

arXiv:2608.09024v1 Announce Type: new Abstract: Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prior

Emotion in an active inference model of human driving

ResearchDGX agent

arXiv:2608.07480v1 Announce Type: new Abstract: Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty reduction. It h

EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations

ResearchDGX agent

arXiv:2608.07497v1 Announce Type: cross Abstract: Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or power

ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document type…

Model ReleasesDGX agent

ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document types, spanning 8 real-world domains: finance, energy, gov, auto

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

Model ReleasesDGX agent

arXiv:2608.09474v1 Announce Type: new Abstract: Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained prima

Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective

SafetyDGX agent

arXiv:2608.08445v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a mechanism to gr

From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing

Model ReleasesDGX agent

arXiv:2608.09842v1 Announce Type: new Abstract: Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on comp

From @sequoia's own your intelligence event. We're starting to see a new category of company emerge, the Full-Stack AI company, where innova…

SafetyDGX agent

From @sequoia's own your intelligence event. We're starting to see a new category of company emerge, the Full-Stack AI company, where innovation happens at both the product and intelligence layer. Tha

Goal-oriented Navigation Instruction Generation with Tour Video Priors

Model ReleasesDGX agent

arXiv:2608.08596v1 Announce Type: new Abstract: Navigation Instruction Generation (NIG) aims to produce step-by-step natural language instructions for navigation guidance. Existing studies primarily t

How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC

ApplicationsDGX agent

ONESTRUCTION, with technical advisory from the AWS Generative AI Innovation Center, built Ishigaki-IDS, a foundation model specialized for construction and BIM workflows. This architectural case study

I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

Model ReleasesDGX agent

Everything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In

I ran Muse Glimmer @ 1M context - All tests passed.

Model ReleasesDGX agent

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Mu

It's been exciting for Ollama to partner with @JensenHuang and the @NVIDIAAI team on launching open models. Open models have no boundaries, …

Model ReleasesDGX agent

It's been exciting for Ollama to partner with @JensenHuang and the @NVIDIAAI team on launching open models. Open models have no boundaries, and let's continue to work together to make this ecosystem b

Learning Multi-Timescale Interventions under Safety and Resource Constraints

SafetyDGX agent

arXiv:2508.03875v2 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas oth

Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

SafetyDGX agent

arXiv:2608.09507v1 Announce Type: cross Abstract: Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain in

LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

Model ReleasesDGX agent

arXiv:2607.15509v2 Announce Type: replace-cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture desig

LogicIF: Towards Complex Logic Instruction Following

Model ReleasesDGX agent

arXiv:2508.09125v4 Announce Type: replace Abstract: Instruction following has catalyzed the recent era of Large Language Models (LLMs) and is the foundational skill underpinning more advanced capabili

MaxModShift: Model Privacy via Designed Shifts

TutorialsDGX agent

arXiv:2608.09328v1 Announce Type: new Abstract: Model learning by an eavesdropper is treated as an estimation problem in a federated environment. The Fisher Information Matrix for the eavesdropper's e

Motif 3: Technical Report

SafetyDGX agent

arXiv:2608.09119v1 Announce Type: new Abstract: We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each spar

Multi-tier storage rewrites the economics of AI inference

HardwareDGX agent

As inference becomes the dominant workload in AI infrastructure, multi-tier storage architectures are emerging as a key method for cost control and enhanced performance. These architectures combine fl

My conversation with @ericvishria of Benchmark. Eric has spent a decade investing across software and hardware, backing companies like Firew…

Model ReleasesDGX agent

My conversation with @ericvishria of Benchmark. Eric has spent a decade investing across software and hardware, backing companies like Fireworks, Sierra, Sunday Robotics, and Cerebras. This one is abo

Nvidia Nemo Switchyard

HardwareDGX agent

https://github.com/NVIDIA-NeMo/Switchyard Finally an open source LLM router. An alternative to openrouter fusion and Sakana Fugu. Doesn't look like it does exactly what Sakana Fugu does according to i

P^{3}: Joint Program-and-Proof Planning for Verified Code Generation

Model ReleasesDGX agent

arXiv:2608.09277v1 Announce Type: new Abstract: Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the program meets a

Planning/RL for a stochastic single-player merge puzzle: afterstates, previewed chance events, and long-horizon throughput [D]

SafetyDGX agent

I am working on an AI for a small single-player merge puzzle and would appreciate pointers to related algorithms, papers, or existing implementations. It resembles 2048 in its action -> afterstate ->

Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction

Local AiDGX agent

arXiv:2608.08443v1 Announce Type: cross Abstract: Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other parts of a

SAIN: Structure-Aware Interactive Navigation with Active Dialogue Grounding for Mobile Robot

Model ReleasesDGX agent

arXiv:2608.09196v1 Announce Type: new Abstract: Most existing vision-language navigation tasks assume that instructions are complete and unambiguous. However, real-world robots often encounter natural

Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access

Model ReleasesDGX agent

arXiv:2608.08942v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide

Sources: Nvidia is developing a Nemotron 4 model with 1T+ parameters, up from Nemotron 3 Ultra's 550B parameters but smaller than leading Chinese open models (The Information)

Model ReleasesDGX agent

The Information: Sources: Nvidia is developing a Nemotron 4 model with 1T+ parameters, up from Nemotron 3 Ultra's 550B parameters but smaller than leading Chinese open models — Nvidia is doubling down

Stealing Reasoning Traces from Proprietary LLM APIs

SafetyDGX agent

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

Model ReleasesDGX agent

arXiv:2608.09538v1 Announce Type: cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.

Test-Time Scaling for CAD Generation via Verifier-Free Consensus Selection

ResearchDGX agent

arXiv:2608.09706v1 Announce Type: cross Abstract: Large language models can write parametric CAD programs from a natural-language description (text-to-CAD generation), but a single sample is often wro

← Previous
1…260261262263264…296
Next →