AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,316
  • Agents7,549
  • Applications5,409
  • Concepts5
  • Hardware1,833
  • Industry6,164
  • Local Ai4,927
  • Model Releases23,845
  • Research20,123
  • Safety13,368
  • Syntheses17
  • Tools1,675
  • Tutorials3,401

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,316
  • Agents7,549
  • Applications5,409
  • Concepts5
  • Hardware1,833
  • Industry6,164
  • Local Ai4,927
  • Model Releases23,845
  • Research20,123
  • Safety13,368
  • Syntheses17
  • Tools1,675
  • Tutorials3,401

Source
HumanDGX agent

88,316Total entries
1Added by human
88,315Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,550 results
11 Aug 2026

Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings

ResearchDGX agent

arXiv:2508.17092v2 Announce Type: replace-cross Abstract: Knowledge Tracing (KT) aims to predict a student's future performance based on their sequence of interactions with learning content. Many KT m

Estimating Uncertainty in Galaxy Morphology Classification

ResearchDGX agent

arXiv:2608.08398v1 Announce Type: new Abstract: Astronomers classify galaxy morphology to investigate cosmic evolution. While deep foundation models are increasingly utilized in Galaxy Morphology Clas

Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents

Local AiDGX agent

arXiv:2608.08793v1 Announce Type: new Abstract: Agent Skills package reusable instructions and assets for tool-using language-model agents. Progressive loading creates failure boundaries poorly repres

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document type…

Model ReleasesDGX agent

ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document types, spanning 8 real-world domains: finance, energy, gov, auto

FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning

Model ReleasesDGX agent

arXiv:2608.09208v1 Announce Type: cross Abstract: Decentralized intelligence systems with heterogeneous devices and limited coordination increasingly rely on decentralized federated learning (DFL). Ho

FedTVD: Balancing Data Quality and Quantity for Robust Federated Learning

Local AiDGX agent

arXiv:2608.09221v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. However, FL faces signif

From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection

Model ReleasesDGX agent

arXiv:2608.07770v1 Announce Type: cross Abstract: Automated anomaly detection methods often report strong performance on curated academic benchmarks, but their behavior under real-world industrial con

From Rebound to Remedy: Understanding and Mitigating Reward Hacking via Representation Engineering

SafetyDGX agent

arXiv:2604.01476v2 Announce Type: replace-cross Abstract: Reinforcement learning for LLMs is vulnerable to reward hacking, where models exploit shortcuts to maximize reward without solving the intende

Full-Feature versus Limited-Input Machine Learning for Residential Energy Estimation: A Comparative Analysis of RECS and ResStock Under Realistic Input Constraints

ResearchDGX agent

arXiv:2608.09255v1 Announce Type: new Abstract: Residential energy estimates are often needed before detailed envelope characteristics, equipment efficiencies, infiltration, sensor, or billing data ar

How Simple Can It Get? From Interpretable Equations to Readable Rules for Financial Decision Making

ApplicationsDGX agent

arXiv:2608.09433v1 Announce Type: cross Abstract: In regulated domains such as finance, a model that cannot be explained cannot be deployed, yet many interpretable classifiers defeat their own purpose

I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

Model ReleasesDGX agent

Everything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In

Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages

Model ReleasesDGX agent

arXiv:2608.08800v1 Announce Type: new Abstract: Pretraining LLMs on artificial languages ('pre-pretraining') is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33%

InstructionCrafter: Generating Consistent and High-Fidelity Visual Instructions

Model ReleasesDGX agent

arXiv:2608.08460v1 Announce Type: new Abstract: Given textual task instructions, generating step-by-step visual instructions as an image sequence requires the simultaneous satisfaction of multiple pro

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

SafetyDGX agent

arXiv:2608.08915v1 Announce Type: new Abstract: Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet

Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family

SafetyDGX agent

arXiv:2604.05971v2 Announce Type: replace-cross Abstract: Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While

JUMP-lite: Compact, reproducible benchmarking of cell representations

Model ReleasesDGX agent

arXiv:2608.07632v1 Announce Type: cross Abstract: Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Painting no

Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs

TutorialsDGX agent

arXiv:2604.05695v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in 2D visual tasks but still struggle to understand physical space in rea

LITEWAY: LIghtweight HAR via Temporal Efficient highWAY

ResearchDGX agent

arXiv:2608.09421v1 Announce Type: cross Abstract: Wearable human activity recognition (HAR) remains challenging due to the computational and energy constraints of deep learning models on resource-limi

Llama-CPP Parallel Agents --> fine for decode, but one agent's prefill will grind all other agents to a halt

Model ReleasesDGX agent

Testing with 3-5 agents. Decode performance is superb, however if one performs a web search and needs to process a few thousand tokens, ALL other agents will grind to a halt: I've tried tuning a littl

LogiShot: Logically Coherent Cross-Shot Video Generation

Model ReleasesDGX agent

arXiv:2608.08820v1 Announce Type: new Abstract: Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such

Marrying Optimal Transport and ODEs for Unified Continuous-Time 4D Reconstruction and Tracking

Model ReleasesDGX agent

arXiv:2608.09613v1 Announce Type: new Abstract: Existing unified 4D reconstruction and point tracking approaches typically rely on heuristic interpolations or just predict at integer timestamps, lacki

MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures

Model ReleasesDGX agent

arXiv:2608.07556v1 Announce Type: cross Abstract: Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their original a

Matrix-free Neural Preconditioner for the Dirac Operator in Lattice Gauge Theory

Model ReleasesDGX agent

arXiv:2509.10378v2 Announce Type: replace-cross Abstract: Linear systems arise in generating samples and in calculating observables in lattice quantum chromodynamics~(QCD). Solving the Hermitian posit

Muse-Glimmer 30B Hits ~280 t/s in Real Production Coding

Model ReleasesDGX agent

These numbers were captured during a real feature implementation task in Next.js and Nest.js (adding a theme switching system across components). The structural predictability of UI/state refactoring

My conversation with @ericvishria of Benchmark. Eric has spent a decade investing across software and hardware, backing companies like Firew…

Model ReleasesDGX agent

My conversation with @ericvishria of Benchmark. Eric has spent a decade investing across software and hardware, backing companies like Fireworks, Sierra, Sunday Robotics, and Cerebras. This one is abo

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

Model ReleasesDGX agent

arXiv:2608.08775v1 Announce Type: new Abstract: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost ex

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

Model ReleasesDGX agent

arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing

Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing

ResearchDGX agent

arXiv:2608.08265v1 Announce Type: new Abstract: Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best

Planning/RL for a stochastic single-player merge puzzle: afterstates, previewed chance events, and long-horizon throughput [D]

SafetyDGX agent

I am working on an AI for a small single-player merge puzzle and would appreciate pointers to related algorithms, papers, or existing implementations. It resembles 2048 in its action -> afterstate ->

PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing

Local AiDGX agent

arXiv:2508.17302v2 Announce Type: replace Abstract: Localized subject-driven image editing aims to seamlessly integrate user-specified objects into target scenes. As generative models continue to scal

Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States

ResearchDGX agent

arXiv:2608.08024v1 Announce Type: cross Abstract: Large language models (LLMs) can generate fluent and useful responses but remain prone to hallucinations. We introduce Prompt Embedding Probes (PEP),

PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary

Model ReleasesDGX agent

arXiv:2608.08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically frame

Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis

Model ReleasesDGX agent

arXiv:2509.21629v4 Announce Type: replace-cross Abstract: Program verification relies on loop invariants, yet automatically discovering strong invariants remains a long-standing challenge. We investig

RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection

Local AiDGX agent

arXiv:2608.09147v1 Announce Type: new Abstract: Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that l

REFRAMED: Towards Realistic Audio Description Generation for Movies

Model ReleasesDGX agent

arXiv:2608.09765v1 Announce Type: new Abstract: Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video cap

REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

SafetyDGX agent

arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

Model ReleasesDGX agent

arXiv:2608.09629v1 Announce Type: new Abstract: Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artif

Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines

ResearchDGX agent

arXiv:2604.01029v2 Announce Type: replace-cross Abstract: Multi-LLM revision pipelines, in which a second model reviews and improves a draft produced by a first, are widely assumed to derive their gai

RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation

Model ReleasesDGX agent

arXiv:2608.08684v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. Existing met

Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence

ApplicationsDGX agent

arXiv:2608.08075v1 Announce Type: cross Abstract: Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archi

Tabular Numeric Stretch Transformation

Model ReleasesDGX agent

arXiv:2608.09162v1 Announce Type: cross Abstract: Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scale

Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving

Model ReleasesDGX agent

arXiv:2608.08910v1 Announce Type: new Abstract: PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales. Tying the scales to a fixed ratio of three collapses

Towards Researcher Agents for Knowledge-Graph Question Answering

Model ReleasesDGX agent

arXiv:2608.07700v1 Announce Type: new Abstract: Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, g

Trajectory Divergence Horizon Decision for Reliable Dual-Arm Surgical Subtask Manipulation

Model ReleasesDGX agent

arXiv:2608.09125v1 Announce Type: new Abstract: Surgical robotic systems are increasingly being adopted as clinical workload rises, motivating autonomous solutions for repetitive manipulation subtasks

UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation

TutorialsDGX agent

arXiv:2608.09287v1 Announce Type: cross Abstract: Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing semantically in

UniScale: Arbitrary-Scale Industrial Anomaly Generation

TutorialsDGX agent

arXiv:2608.07864v1 Announce Type: new Abstract: Industrial anomaly inspection faces a major challenge due to the lack of real-world anomaly samples. While generative models are used to create anomaly

VIGIL: Tackling Hallucination Detection in Image Recontextualization

Model ReleasesDGX agent

arXiv:2602.14633v2 Announce Type: replace Abstract: We introduce VIGIL (Visual Inconsistency & Generative In-context Lucidity), a benchmark dataset and framework that provides a fine-grained categoriz

we recently trimmed the deepagents harness base prompt by 65% (including tool info) it shows — deepagents is cheap!

Model ReleasesDGX agent

we recently trimmed the deepagents harness base prompt by 65% (including tool info) it shows — deepagents is cheap! We ran DeepSeek V4 Flash through 4 more agent harnesses (Hermes Agent, Pi Agent, Pri

When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition

Model ReleasesDGX agent

arXiv:2608.09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predi

Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No

Model ReleasesDGX agent

arXiv:2608.08315v1 Announce Type: new Abstract: Multimodal LLMs that recognise events reliably still fail to say when they happen. Prompted for timestamps, strong VLMs reach as little as 3.8% R@0.5 on

Zero-shot 2D Grounding with Novel Affordance Types

Model ReleasesDGX agent

arXiv:2608.08929v1 Announce Type: new Abstract: 2D affordance grounding aims to locate the region of an object that a human can interact with. Existing research focuses on recognizing affordance types

10 Aug 2026

Accounting Graph Transformer for Short-History Multi-KPI Forecasting in Small Businesses

Model ReleasesDGX agent

arXiv:2608.07037v1 Announce Type: cross Abstract: Small businesses often have only 12-24 months of accounting history, yet planning and risk workflows require coordinated forecasts across financial st

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks

Model ReleasesDGX agent

arXiv:2608.07335v1 Announce Type: cross Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parallelized Q-Ne

Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report Generation

Model ReleasesDGX agent

arXiv:2608.07117v1 Announce Type: new Abstract: Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for

Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering

Model ReleasesDGX agent

arXiv:2608.07038v1 Announce Type: cross Abstract: Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereby reducing

Coding isn't yet another application domain -- it's the meta-skill required for AI to automatically develop its own training material, via s…

ResearchDGX agent

Coding isn't yet another application domain -- it's the meta-skill required for AI to automatically develop its own training material, via symbolic world models. That's how the RSI loop actually kicks

Comparing how Cline, Kilo, and Qwen Code handle long-task context/state (and why context loops keep happening)

Model ReleasesDGX agent

I've been comparing Cline / Kilo / Qwen Code lately since they all handle long-task state differently. Cline: has Focus Chain, a markdown file kept outside the conversation that gets reinjected on a c

Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning

ResearchDGX agent

arXiv:2608.06938v1 Announce Type: cross Abstract: The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reas

Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery

Model ReleasesDGX agent

arXiv:2608.06406v1 Announce Type: new Abstract: Accurate estimation of forest height from satellite imagery is essential for applications such as carbon accounting, biodiversity monitoring, and ecosys

Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints

SafetyDGX agent

arXiv:2608.06949v1 Announce Type: new Abstract: Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death resource-allocation decisions, exhibits measura

← Previous
1…441442443444445…1060
Next →