AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,343
  • Agents7,552
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,164
  • Local Ai4,928
  • Model Releases23,861
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,402

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,343
  • Agents7,552
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,164
  • Local Ai4,928
  • Model Releases23,861
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,402

Source
HumanDGX agent

88,343Total entries
1Added by human
88,342Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,572 results
20 Aug 2026

The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches

Model ReleasesDGX agent

Component Validated configuration Motherboard ASRock Rack SPC621D8U-2T/OVH CPU Xeon Gold 6330 (Get gold/platinum if interested in Optane Pmem gimmicks) GPU fabric Two Broadcom/PLX PEX88096 islands, ei

The real benchmark for an AI agent isn’t what it says. It’s what it gets done. AutoClaw, @Zai_org 's AI Agent for Work, is built to turn int…

Model ReleasesDGX agent

The real benchmark for an AI agent isn’t what it says. It’s what it gets done. AutoClaw, @Zai_org 's AI Agent for Work, is built to turn intent into finished output. Give it a complex task. AutoClaw b

The real lesson here is that if you pick and choose your stats (and which quarters to show, and so on) you can tell any story you like.

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

The real lesson here is that if you pick and choose your stats (and which quarters to show, and so on) you can tell any story you like. May startle the markets with this but: OpenAI is growing faster

TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion

Model ReleasesDGX agent

arXiv:2601.17178v3 Announce Type: replace-cross Abstract: Hardware Trojans (HTs) remain a critical threat because learning-based detectors often overfit to narrow trigger/payload patterns and small, s

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

Model ReleasesDGX agent

arXiv:2608.18504v1 Announce Type: new Abstract: Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching and fine-graine

Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis

ResearchDGX agent

arXiv:2608.18825v1 Announce Type: cross Abstract: Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilingual use ca

What Makes Software Issue Resolution Tasks Difficult for Agents?

Model ReleasesDGX agent

arXiv:2608.18280v1 Announce Type: cross Abstract: Background. Advances in agentic systems are simultaneously, and rapidly, saturating benchmarks. Despite this often discussed phenomena, benchmark scor

When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift

ResearchDGX agent

arXiv:2608.18330v1 Announce Type: new Abstract: Whether input-dependent ('dynamic') combination of a regression model pool beats the best static blend depends on the shift and is rarely known before d

When Two Tracers Disagree: An Investigation of Multimodal Fusion for Clinical PET/CT Segmentation

ResearchDGX agent

arXiv:2608.19063v1 Announce Type: new Abstract: PSMA and FDG PET/CT visualise complementary biological information in prostate cancer. Combining both tracers could capture heterogeneous tumour phenoty

19 Aug 2026

4 days to the top.🏆Thank you to every builder who pushed Qwen3.8-27B to #1 on Cline! @cline

Local AiDGX agent

4 days to the top.🏆Thank you to every builder who pushed Qwen3.8-27B to #1 on Cline! @cline Qwen3.8-27B is now the #1 local model in Cline after just 4 days. This ends a *4 month* streak by the previo

A Diffusion-Refined Planner with Reinforcement Learning Priors for Confined-Space Parking

SafetyDGX agent

arXiv:2510.14000v2 Announce Type: replace Abstract: The growing demand for parking has increased the need for automated parking planning methods that can operate reliably in confined spaces. In restri

An Investigation of the NeurIPS and ICML 2025 Position Tracks

Model ReleasesDGX agent

arXiv:2608.16894v1 Announce Type: cross Abstract: ML venues shape what kinds of research claims become legible to reviewers and what forms of evidence count as rigorous. The NeurIPS and ICML Position

ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generation

Model ReleasesDGX agent

arXiv:2608.17356v1 Announce Type: new Abstract: Most automated essay scoring (AES) systems output a single holistic score without interpretable evidence and rely on closed APIs that introduce data pri

Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch

Model ReleasesDGX agent

arXiv:2608.17684v1 Announce Type: new Abstract: Self-evolving agents turn experience into reusable skills, workflows, or memories, but post-evolution accuracy alone does not show whether learned behav

b10502

Model ReleasesDGX agent

ci : add attestation for signed release artifacts (#25933) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/41614541 macOS/iOS: macOS Apple Silicon (arm64) m

Benchmarking Automated Security Patch Backporting: How Far Are We?

Model ReleasesDGX agent

arXiv:2608.17671v1 Announce Type: cross Abstract: Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their respective

CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion

Model ReleasesDGX agent

arXiv:2608.17911v1 Announce Type: new Abstract: As LLM agents operate across structured workflows and sessions, preserving long-term history does not ensure that later contexts can recover relevant ev

Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints

AgentsDGX agent

arXiv:2608.17932v1 Announce Type: cross Abstract: Groups routinely complete projects that no single member can plan, execute, or verify alone. We propose a formal model of this phenomenon, Collective

Comprehensive framework for evaluation of deep neural networks in detection and quantification of lymphoma from PET/CT images: clinical insights, pitfalls, and observer agreement analyses

ResearchDGX agent

arXiv:2311.09614v5 Announce Type: replace-cross Abstract: This study addresses critical gaps in automated lymphoma segmentation from PET/CT images, focusing on issues often overlooked in existing lite

Conceptual integrity and counting lines of code

Model ReleasesDGX agent

Last week I recorded an episode of the Talking Postgres podcast with Claire Giordano on the subject of 'How AI is changing software development'. We had a really great conversation. Here are a couple

Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation

Model ReleasesDGX agent

arXiv:2608.18034v1 Announce Type: new Abstract: Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coh

Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating

AgentsDGX agent

arXiv:2608.18058v1 Announce Type: new Abstract: Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet their viability depends on a condition

Detecting and Discriminating Operator Misspecification in Hybrid PDE-Parameter Learning: a Reference-Free Instrument, with Discrimination Bounded In Sample

Model ReleasesDGX agent

arXiv:2608.16925v1 Announce Type: new Abstract: We build an instrument that reads, from a single fit and with no oracle, whether the operator a hybrid PDE-parameter estimator postulates is wrong-and s

Efficient Dynamic Shielding for Parametric Safety Specifications

Model ReleasesDGX agent

arXiv:2505.22104v2 Announce Type: replace Abstract: Shielding has emerged as a promising approach for ensuring safety of AI-controlled autonomous systems. The algorithmic goal is to compute a shield,

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

SafetyDGX agent

arXiv:2608.17512v1 Announce Type: new Abstract: Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing m

Eval4Sim: An Evaluation Framework for Persona Simulation

SafetyDGX agent

arXiv:2603.02876v2 Announce Type: replace Abstract: Large Language Model personas, explicit profiles specifying a user's attributes, preferences, and behavioural tendencies, are increasingly used to s

Exact Reformulation and Optimization for Direct Metric Optimization in Binary Imbalanced Classification

Model ReleasesDGX agent

arXiv:2507.15240v2 Announce Type: replace Abstract: For classification with imbalanced class frequencies, i.e., imbalanced classification (IC), standard accuracy is known to be misleading as a perform

Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification

Model ReleasesDGX agent

arXiv:2608.17247v1 Announce Type: new Abstract: Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this

FedPref: Federated Preference Learning for Structured Radiology Report Extraction

Local AiDGX agent

arXiv:2608.16971v1 Announce Type: new Abstract: Radiology reports describe findings and locations in free text, but downstream search and analysis require these relations in a fixed schema. Learning t

FetchMan: Learning Visual Humanoid Loco-Manipulation Policies from Simulated Experiences

Model ReleasesDGX agent

arXiv:2608.17027v1 Announce Type: new Abstract: Visual loco-manipulation policies that can generalize to novel scenes and objects have long been a goal of robotics research. However, today's data-hung

From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

ResearchDGX agent

arXiv:2608.17827v1 Announce Type: new Abstract: Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as

Google announces new study tools, including a student hub, notebooks, and interactive 3D visualizations in Gemini, plus student offers for Google AI plans (Amanda Caswell/Tom's Guide)

Model ReleasesDGX agent

Amanda Caswell / Tom's Guide: Google announces new study tools, including a student hub, notebooks, and interactive 3D visualizations in Gemini, plus student offers for Google AI plans — Head back to

Grok 4.6 is #1 on healthcare questions

Model ReleasesDGX agent

Grok 4.6 is #1 on healthcare questions Grok 4.6 just took the #1 spot on MedAgentBench....one of the most interesting benchmarks for real-world agentic healthcare tasks and Grok now has two generation

GSToken: Geometry-Structured Gaussian Tokens for Compact 3D Medical Image Representation

Model ReleasesDGX agent

arXiv:2608.17425v1 Announce Type: new Abstract: Effective segmentation of multi-modal MRI is central to improving neural network accuracy in brain tumor recognition. Existing methods typically compres

Harness launches AI agents that triage and patch vulnerabilities

Model ReleasesDGX agent

Software delivery platform provider Harness Inc. today launched a set of artificial intelligence agents that find software vulnerabilities and write the patches. Developers approve the fixes before an

hot take: you shouldn’t have to know whether an idea is worth building before you get to try it 😅 I find that when every build has a cost, …

Model ReleasesDGX agent

hot take: you shouldn’t have to know whether an idea is worth building before you get to try it 😅 I find that when every build has a cost, you start judging your ideas before you know what they could

How smoothing the affinity matrix affects neighborhood preservation in t-SNE

Model ReleasesDGX agent

arXiv:2608.17190v1 Announce Type: cross Abstract: Dimensionality reduction methods are instrumental to visualize high-dimensional data, and t-SNE stands as one of the most widely used methods due to i

I think folks are going to be stunned by how quickly Grok starts generating massive revenue for SpaceX. Grok Bot is easily one of the most u…

Model ReleasesDGX agent

I think folks are going to be stunned by how quickly Grok starts generating massive revenue for SpaceX. Grok Bot is easily one of the most useful (and addicting) AI products out there. It’s very Apple

Initialization-Free Bundle Adjustment Revisited: A Controlled Experimental Study

Model ReleasesDGX agent

arXiv:2608.18028v1 Announce Type: new Abstract: Initialization-free bundle adjustment (InitFree BA) aims to recover camera poses and scene structure directly from image observations, avoiding the geom

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

Model ReleasesDGX agent

arXiv:2608.17071v1 Announce Type: new Abstract: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in

Lymphocyte Mimicry Correction via Region-Level Tissue Reasoning and Unbalanced Optimal Transport

ResearchDGX agent

arXiv:2608.17151v1 Announce Type: new Abstract: Cell mimicry arises when different cell types appear morphologically similar. Human pathologists resolve this ambiguity using surrounding tissue context

MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps

Model ReleasesDGX agent

arXiv:2608.17659v1 Announce Type: cross Abstract: LLM-powered GUI agents that autonomously operate smartphones are rapidly transitioning from research prototypes to early real-world deployment. Howeve

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

ResearchDGX agent

arXiv:2608.17402v1 Announce Type: new Abstract: Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling

Most founders copy/paste their privacy policy from Stripe or ask Claude to draft an MSA at 2:00 AM, then pray they don't get sued. You can s…

Model ReleasesDGX agent

Most founders copy/paste their privacy policy from Stripe or ask Claude to draft an MSA at 2:00 AM, then pray they don't get sued. You can stop doing this. We built a free, CC0-licensed legal template

NeuroPath: Brain-Inspired Dual-Pathway Graph Convolutional Networks for Skeleton-Based Action Recognition

ResearchDGX agent

arXiv:2608.17487v1 Announce Type: new Abstract: Skeleton-based action recognition aims to recognize human actions from sequences of human joint coordinates. Most existing Spatial-Temporal Graph Convol

Online Generalized Sparse Regression: How Does Overparametrization Help?

Model ReleasesDGX agent

arXiv:2608.17466v1 Announce Type: cross Abstract: Regularized sparse regression has been extensively studied in the offline setting, but online formulation remains relatively under-explored. This gap

Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization

ResearchDGX agent

arXiv:2608.18040v1 Announce Type: cross Abstract: Sampling from a diffusion model typically requires many forward passes through a large neural network, making generation computationally expensive. Wh

OV3D-Bench: A Diagnostic Benchmark for Open-Vocabulary Monocular 3D Detection

Model ReleasesDGX agent

arXiv:2608.17110v1 Announce Type: new Abstract: Open-vocabulary monocular 3D detectors report strong in-domain performance, but each evaluates under a different protocol, several rely on per-image cat

Parametric Knowledge in RAG-SFT for Domain-Specific Document Generation

ResearchDGX agent

arXiv:2603.23047v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) fine-tuning has shown substantial improvements over vanilla RAG, yet most studies target document questio

Picture the Epsilon: Pursuing Identity-Level Privacy Guarantees for Images

Model ReleasesDGX agent

arXiv:2608.17147v1 Announce Type: cross Abstract: Image-to-image face generators are widely used, and visual dissimilarity between their outputs and source images is sometimes treated as evidence of p

PRISM: Precision and contact-rich Real-world Industrial Skill dataset with Multimodal sensing

Model ReleasesDGX agent

arXiv:2608.17962v1 Announce Type: new Abstract: Recent progress in robotic learning has been fueled by large-scale datasets collected in everyday environments. However, most existing datasets emphasiz

PrivAct: Internalizing Contextual Privacy Preservation via Multi-Agent Preference Training

AgentsDGX agent

arXiv:2602.13840v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly deployed in personalized tasks involving sensitive, context-dependent information, where privacy

Procedural Content Metageneration via Program Search and Continual Abstraction Discovery

ResearchDGX agent

arXiv:2608.17947v1 Announce Type: new Abstract: Large language models can generate executable programs, which makes it possible to search directly over procedural content generators rather than indivi

QuantumNovelty: A Skill-Orchestrating Language Agent for Referee-Style Review and Patentability Screening of Quantum Papers and Patents

AgentsDGX agent

arXiv:2608.16900v1 Announce Type: cross Abstract: Language-model agents increasingly produce quantum-science results; we ask whether the same agentic paradigm can also scrutinize them in an auditable,

Qwen 3.8 27B SlopCodeBench results

Model ReleasesDGX agent

Howdy, I'm back again - running my favorite benchmark (it's still unsaturated for the time being so might as well!) previous runs a b https://github.com/michaelasper/benchmarks/blob/main/qwen3.8-27b-p

Reinforcing Consistency in Video MLLMs with Structured Rewards

SafetyDGX agent

arXiv:2604.01460v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding. However, seemingly plausible outputs often suffer

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

ResearchDGX agent

arXiv:2608.17426v1 Announce Type: cross Abstract: We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achie

smolmachines / smolvm as a sandbox for untrusted Python & JavaScript

Model ReleasesDGX agent

Research: smolmachines / smolvm as a sandbox for untrusted Python & JavaScript I tasked Claude Fable 5 running in Claude Code for web with the following research task: Put https://smolmachines.com thr

Supporting Calibrated Reliance in Human-AI Collaboration: Different Strategies for Different Tasks

Model ReleasesDGX agent

arXiv:2604.03237v2 Announce Type: replace-cross Abstract: As AI systems increasingly support human decision making, a central challenge is determining what information helps people recognize when to r

The submissions for the $2M Build with Gemini XPRIZE just closed. Over 90 days 26k competitors built a real AI-native business solving a pro…

Model ReleasesDGX agent

The submissions for the $2M Build with Gemini XPRIZE just closed. Over 90 days 26k competitors built a real AI-native business solving a problem in the market. No demos allowed, only real users with r

← Previous
1…592593594595596…1060
Next →