AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
Human
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
83,192 results
12 Aug 2026

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

Model ReleasesDGX agent

arXiv:2608.10042v1 Announce Type: cross Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool u

UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference

Model ReleasesDGX agent

arXiv:2603.18446v2 Announce Type: replace Abstract: Long-context inference remains challenging for large language models due to attention dilution and out-of-distribution degradation. Context selectio

V-FiLLM: Verified Financial LLM Reasoning Benchmark

Model ReleasesDGX agent

arXiv:2608.11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains compar

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

v0.32.10

Model ReleasesDGX agent

What's Changed Models that don't set a repeat_penalty now default to 1.0 (off) instead of 1.1, matching other engines and speeding up speculative decoding; set a per-model parameter if an older model

v0.32.10-rc0: nn: speed up prefill on double-scale nvfp4 models

Local AiDGX agent

ModelOpt checkpoints apply a float32 global scale to every projection output on top of the per-group quantization scales. Running the multiply and the cast back to the activation dtype as separate eag

Validated Synthetic Patient Generation for Small Longitudinal Cohorts: Coagulation Dynamics Across Pregnancy

ResearchDGX agent

arXiv:2604.07557v2 Announce Type: replace Abstract: Small longitudinal cohorts, common in maternal health, rare diseases, and early-phase trials, limit computational modeling because enrollment is slo

VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection

AgentsDGX agent

arXiv:2511.19436v2 Announce Type: replace-cross Abstract: Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary mode

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus

ResearchDGX agent

arXiv:2608.10665v1 Announce Type: new Abstract: Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approache

Very interesting new work from Anthropic. (bookmark it) They evolve mind viruses, ideas that spread through a multi-agent system by getting …

AgentsDGX agent

Very interesting new work from Anthropic. (bookmark it) They evolve mind viruses, ideas that spread through a multi-agent system by getting each host to pass them on, then measure what governs the spr

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

Model ReleasesDGX agent

arXiv:2608.10875v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained re

VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

Local AiDGX agent

arXiv:2608.11201v1 Announce Type: new Abstract: Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and auth

VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation

SafetyDGX agent

arXiv:2608.10903v1 Announce Type: new Abstract: Reliable clinical deployment of machine learning requires models that know when they are likely to fail, particularly for subgroups underrepresented in

VIScore: Diagnosing Planning-Relevant Quality in Latent World Models

ResearchDGX agent

arXiv:2608.11174v1 Announce Type: new Abstract: Regulating the latent space to an isotropic Gaussian distribution provides a stable and information-maximized landscape for world model planning. Howeve

VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

Model ReleasesDGX agent

arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world

Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps

Model ReleasesDGX agent

arXiv:2607.16173v2 Announce Type: replace Abstract: Open-vocabulary 3D maps let robots answer language queries about what and where, but they assume a static world and cannot answer queries about how

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

Model ReleasesDGX agent

arXiv:2608.10682v1 Announce Type: new Abstract: Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing

VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation

Model ReleasesDGX agent

arXiv:2608.10359v1 Announce Type: cross Abstract: As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, l

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning

SafetyDGX agent

arXiv:2608.11013v1 Announce Type: new Abstract: Text-only training is a popular paradigm in zero-shot video captioning, where the video distribution is not available to the model during training, lead

WaveInst: A Frequency-Domain Enhanced Network for Fine-Grained Thin Tree Trunk Extraction in Forest Scenes

ResearchDGX agent

arXiv:2505.01656v2 Announce Type: replace Abstract: Analyzing tree morphology, particularly trunk and branch extraction, is valuable for genetic breeding and forestry management. Existing image-based

We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: …

Model ReleasesDGX agent

We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B

We built a software factory for AI SDK. Each step is an agent, and humans merge changes. Four weeks in: ▪️ The factory authors up to 35% of …

AgentsDGX agent

We built a software factory for AI SDK. Each step is an agent, and humans merge changes. Four weeks in: ▪️ The factory authors up to 35% of merged PRs ▪️ It closed 70% of issues in July ▪️ Open bugs a

We have her dash camera which shows they are lying. The agents are wearing body cameras and should have dash cameras of their own. If what t…

Model ReleasesDGX agent

We have her dash camera which shows they are lying. The agents are wearing body cameras and should have dash cameras of their own. If what they say happened was true they wouldn’t be issuing statement

We wrote a 36-page ArXiv whitepaper on ExtractBench 🧑‍🔬 , our effort to create the most comprehensive, schema-guided, real-world document …

Model ReleasesDGX agent

We wrote a 36-page ArXiv whitepaper on ExtractBench 🧑‍🔬 , our effort to create the most comprehensive, schema-guided, real-world document extraction benchmark. It’s extremely detailed and covers every

Weighted Sequential Bayesian Inference for Non-Stationary Linear Contextual Bandits

ResearchDGX agent

arXiv:2307.03587v4 Announce Type: replace Abstract: In non-stationary linear contextual bandits, existing efficient algorithms typically rely on the Weighted Regularized Least-Squares (WRLS) estimator

What DINO saw: ALiBi positional encoding reduces positional bias in Vision Transformers

SafetyDGX agent

arXiv:2603.16840v2 Announce Type: replace Abstract: Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However

What do you guys do for GPU Kernels?

HardwareDGX agent

I'm trying to figure out GPU Kernel optimization on older hardware like SM80(ampere) . Is there tools you guys use? Or frameworks? Im waiting for this framework https://www.reddit.com/r/LocalLLaMA/com

What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model

AgentsDGX agent

arXiv:2608.10986v1 Announce Type: new Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such

What unique, custom QOL upgrades have you given your local agents?

Model ReleasesDGX agent

Warning: Kinda long post. If you don't like reading, please skip for your own sanity. Also, I've got nothing to sell, just a tinkerer, so I just want to share ideas and learn from you guys too. When I

What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research

SafetyDGX agent

arXiv:2608.10431v1 Announce Type: cross Abstract: Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpret, implemen

When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting

AgentsDGX agent

arXiv:2606.16465v2 Announce Type: replace Abstract: AI agents can now take irreversible actions in operational systems, but agent-caused losses are still not clearly assigned, priced, or transferred.

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

Model ReleasesDGX agent

arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of

When Do Anchor-Based Pointwise LLM Rerankers Help? Retriever Quality, Statistical Scope, and Anchor Design

ResearchDGX agent

arXiv:2608.10528v1 Announce Type: cross Abstract: Anchor-based pointwise LLM reranking scores each candidate against a shared reference passage to recover cross-document context at pointwise cost. We

When Is a General Factor Distinguishable? Non-Proportionality, Stable Structure, and the Bifactor Decision

ResearchDGX agent

arXiv:2608.10731v1 Announce Type: cross Abstract: Whether an additional general dimension is necessary beyond correlated first-order factors is a property of the population covariance matrix, not of a

When Repository Labels Are Not Image-Level Truth: A Supervision Auditing Framework for Chest Radiograph AI

ApplicationsDGX agent

arXiv:2608.10084v1 Announce Type: cross Abstract: Public chest X-ray repositories are widely used to train medical AI systems, yet their labels are typically extracted from radiology reports rather th

When should you start post-training your own models? @FireworksAI_HQ CEO @lqiao’s answer: after product-market fit. Not because it's hard...…

SafetyDGX agent

When should you start post-training your own models? @FireworksAI_HQ CEO @lqiao’s answer: after product-market fit. Not because it's hard... but because only after PMF is the data coming off your prod

When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs

Local AiDGX agent

arXiv:2608.10489v1 Announce Type: new Abstract: Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token p

When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.11024v1 Announce Type: new Abstract: Attribute hallucination---where vision-language models (VLMs) correctly identify an object but mischaracterize its properties---is prevalent yet mechani

When Your State Estimator Has Lost The Plot: Detecting Estimator Failures Via Spectral Analysis

ApplicationsDGX agent

arXiv:2608.10623v1 Announce Type: new Abstract: Reliable onboard state estimation is essential for safe robotic operation, yet unmodeled disturbances, such as sensor aliasing or out-of-distribution no

Where To Look? : Causal Tracing of Vision Encoders in VLM

ResearchDGX agent

arXiv:2608.10758v1 Announce Type: new Abstract: Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actua

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition

ResearchDGX agent

arXiv:2608.10836v1 Announce Type: cross Abstract: The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory tr

Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking

ApplicationsDGX agent

arXiv:2608.10329v1 Announce Type: cross Abstract: Notice-and-comment rulemaking gives any affected party the same formal right to influence federal regulation, but formal access is not substantive cap

Whole-Body Planning for Humanoids Navigating Confined Spaces via Self-Collision Avoidance References

Model ReleasesDGX agent

arXiv:2608.10220v1 Announce Type: new Abstract: Humanoid locomotion in highly confined environments requires navigating dense environmental obstacles and complex self-collision bounds while maintainin

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

Model ReleasesDGX agent

arXiv:2608.11095v1 Announce Type: new Abstract: Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wh

Why does CLAUDE.MD keep growing? If you maintain a CLAUDE.md or an AGENTS.md, this one is worth your time. (bookmark it) This work traces wh…

Model ReleasesDGX agent

Why does CLAUDE.MD keep growing? If you maintain a CLAUDE.md or an AGENTS.md, this one is worth your time. (bookmark it) This work traces why these files grow without bound. Appending an instruction i

Wind-Informed Rapid Flight-Planning in Complex Urban Topologies via Machine Learning and Experimental Validation

ResearchDGX agent

arXiv:2608.10309v1 Announce Type: cross Abstract: Advanced air mobility operations hold the potential to enhance and expand regional transportation of both people and goods in populated areas. However

With the smartest person I know, @HarshSensei, we're building evsys-sdk for the community, this open-source repository allows anyone to buil…

Model ReleasesDGX agent

With the smartest person I know, @HarshSensei, we're building evsys-sdk for the community, this open-source repository allows anyone to build their own continual learning system with first-class suppo

Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output

Model ReleasesDGX agent

arXiv:2608.10279v1 Announce Type: cross Abstract: Streaming language-model output creates a release-timing problem: complete-response moderation acts after streamed text has escaped, whereas repeated

Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

Model ReleasesDGX agent

arXiv:2608.11022v1 Announce Type: cross Abstract: Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their con

Write your first prompt with the GitHub Copilot app

TutorialsDGX agent

Learn how to write your first prompt in the GitHub Copilot app, choose the right context and model, and start your first task with confidence. The post Write your first prompt with the GitHub Copilot

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

ResearchDGX agent

arXiv:2608.10878v1 Announce Type: new Abstract: Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchanne

XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving

SafetyDGX agent

arXiv:2608.10976v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verb

'YES! YES! I absolutely love this insight!' Affirmative Narration as Interactional Strategy in Dialogues with LLM Chatbots

ResearchDGX agent

arXiv:2607.28646v2 Announce Type: cross Abstract: This article analyses narrative mechanisms that are common in dialogues with LLM chatbots. In combination, these mechanisms produce an interactional s

You can now use Ollama as a provider in GitHub Copilot for JetBrains. https://github.blog/changelog/2026-08-11-copilot-memory-and-ollama-in-…

Local AiDGX agent

GitHub announced on August 12 2026 that users can now integrate Ollama as a provider in **GitHub Copilot for JetBrains**. This update allows JetBrains developers to switch to or add locally‑hosted (or

You chose the best model. Why is your agent still failing?

AgentsDGX agent

Public benchmarks can show how a model performs in general. Production reliability depends on the context and harness around it, which only your team can evaluate against its own data, workflows, and

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

SafetyDGX agent

arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream dec

Yukon: Open Innovation for Frontier Research Open Innovation has been the driving value at Eigen Labs. But I’ve often struggled with where i…

AgentsDGX agent

Yukon: Open Innovation for Frontier Research Open Innovation has been the driving value at Eigen Labs. But I’ve often struggled with where it is actually better than a great closed team. Finally we fo

ZeroPur: Succinct Training-Free Adversarial Purification

ResearchDGX agent

arXiv:2406.03143v4 Announce Type: replace Abstract: Adversarial purification is a kind of defense technique that can defend against various unseen adversarial attacks without modifying the victim clas

11 Aug 2026

1 Day in and I feel okay saying Muse-Glimmer-30B finally beats 3.6-27B for the size in some use-cases

Model ReleasesDGX agent

A few things right off the bat: it reasons very efficiently. Like Grok 4.5 levels of efficient thinking it quantizes very well. My first few tests with iq3_xxs were better than Qwen/Gemma behaved at t

10 year garbage card for local llms

Model ReleasesDGX agent

Hello everyone! ​I like dumb things. I like working with weak computers and microcontrollers. I like the simplicity and low electricity usage. Simply put, the efficiency of a 'dumb' PC. ​The first tim

12GB VRAM gang, what's our plan?

Model ReleasesDGX agent

Seems like we're limited to qwen finetuned MoEs for now. Looking at the current landscape - focus seems to be on dense models (muse glimmer 30b, qwen 3.8 27b) for smaller setups. Is upgrading to 24GB

← Previous
1…1011121314…1387
Next →