AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,612 results
29 May 2026

LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis

Model ReleasesDGX agent

arXiv:2605.28876v1 Announce Type: cross Abstract: CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy. Coding agents that try to debug them depend on an upstream tool to red

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference

Model ReleasesDGX agent

arXiv:2510.24606v2 Announce Type: replace Abstract: The quadratic cost of attention limits the scalability of long-context LLMs, especially under limited hardware memory budgets. While attention is of

Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.30274v1 Announce Type: cross Abstract: Document-level translation remains one of the most challenging tasks for large language models, which are constrained by limited context windows that

LoopFM: Learning frOm HistOrical RePresentations of Foundation Model for Recommendation

Model ReleasesDGX agent

arXiv:2605.29280v1 Announce Type: cross Abstract: Knowledge distillation (KD) transfers a single scalar prediction from a large foundation model (FM) to compact vertical models (VMs), suffering from d

love it! @ggerganov 's llama.cpp delivering!

Model ReleasesDGX agent

love it! @ggerganov 's llama.cpp delivering! pibot is now running fully local, using parakeet for STT, qwen3-tts for TTS, and Qwen 3.6 as the local multi-modal LLM via llama.cpp. The STT and TTS infer

LUMINA: A Multi-Vendor Mammography Benchmark with Energy Harmonization Protocol

Model ReleasesDGX agent

arXiv:2603.14644v3 Announce Type: replace-cross Abstract: Publicly available full-field digital mammography (FFDM) datasets remain limited in size, clinical annotations, and vendor diversity, hinderin

MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

Model ReleasesDGX agent

arXiv:2601.04633v2 Announce Type: replace Abstract: Machine-Generated Text (MGT) is becoming increasingly difficult to distinguish from Human-Written Text (HWT). This trend has exacerbated malicious a

MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference

Model ReleasesDGX agent

arXiv:2605.30218v1 Announce Type: new Abstract: Temperature-zero BF16 LLM inference is often treated as reproducible, yet the same request can emit different tokens when decoded alone or inside a larg

MATNet: Multi-Level Fusion Transformer-Based Model for Day-Ahead PV Generation Forecasting

Model ReleasesDGX agent

arXiv:2306.10356v3 Announce Type: replace-cross Abstract: Accurate forecasting of renewable generation is crucial to facilitate the integration of Renewable Energy Sources into the power system. Focus

Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?

Model ReleasesDGX agent

arXiv:2605.28860v1 Announce Type: cross Abstract: Fine-tuning large language models (LLMs) frequently induces catastrophic forgetting of prior capabilities. Recent work has shown that reinforcement le

MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models

Model ReleasesDGX agent

arXiv:2605.28825v1 Announce Type: new Abstract: Large language models (LLMs) frequently encode factual and reasoning knowledge in their internal representations that is not faithfully reflected in the

MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models

Model ReleasesDGX agent

arXiv:2507.09574v3 Announce Type: replace-cross Abstract: Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requ

Mind Your Tone: Does Tone Alter LLM Performance?

Model ReleasesDGX agent

arXiv:2605.29027v1 Announce Type: new Abstract: The use of Large Language Models (LLMs) is proliferating, yet their performance is observed to vary based on prompting styles and tones. In this study,

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

Model ReleasesDGX agent

arXiv:2605.29737v1 Announce Type: cross Abstract: LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code t

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

Model ReleasesDGX agent

arXiv:2605.29360v1 Announce Type: new Abstract: Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that t

Mitigating Hallucination in Vision-Language Models through Barrier-Regulated Adaptive Closed-form Steering

Model ReleasesDGX agent

arXiv:2605.29881v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) often hallucinate objects that are not present in the input image, largely because visual grounding weakens as de

Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions

Model ReleasesDGX agent

arXiv:2605.29862v1 Announce Type: cross Abstract: AI-driven respiratory sound classification (RSC) is promising for automated pulmonary disease detection, yet multi-site deployment is hindered by inte

models underestimate how much work it takes (token usage) to accomplish a task, just like us

Model ReleasesDGX agent

models underestimate how much work it takes (token usage) to accomplish a task, just like us 🧵 Claude-Opus-4.8 takes you too much tokens - but is this issue general across agents? Do agents know how m

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing

Model ReleasesDGX agent

arXiv:2605.22100v2 Announce Type: replace Abstract: Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information sys

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

Model ReleasesDGX agent

arXiv:2605.29738v1 Announce Type: cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual

Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2602.14399v2 Announce Type: replace Abstract: Multi-turn jailbreak attacks have proven effective against text-only large language models (LLMs), where malicious content is gradually introduced t

Multimodal LLMs See Sentiment

Model ReleasesDGX agent

arXiv:2508.16873v3 Announce Type: replace Abstract: Understanding how visual content conveys sentiment is increasingly important in a digital landscape dominated by imagery. However, sentiment percept

MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization

Model ReleasesDGX agent

arXiv:2605.29951v1 Announce Type: new Abstract: Understanding how harm emerges from interaction between otherwise benign image-text pairs requires intent-aware cross-modal reasoning beyond surface-lev

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

Model ReleasesDGX agent

arXiv:2605.29300v1 Announce Type: cross Abstract: Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses ar

NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs

Model ReleasesDGX agent

arXiv:2605.29716v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising non-autoregressive generative paradigm. Given the prohibitive computational cost of

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

Model ReleasesDGX agent

arXiv:2605.29685v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social

No More K-means:Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval

Model ReleasesDGX agent

arXiv:2605.30120v1 Announce Type: cross Abstract: Multi-vector retrieval (MVR) models, exemplified by ColBERT, have established new benchmarks in retrieval accuracy by preserving fine-grained token-le

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

Model ReleasesDGX agent

arXiv:2605.29676v1 Announce Type: new Abstract: Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default languag

Nvidia up 0.7%, on news that tokenmaxxing is dead and H200 rental prices are down. What an absurd time to be alive.

Model ReleasesDGX agent

Nvidia up 0.7%, on news that tokenmaxxing is dead and H200 rental prices are down. What an absurd time to be alive. In the last 30 days alone: – Microsoft cancelled most of its Claude Code licenses, c

OISD: On-Policy Internal Self-Distillation of Language Models

Model ReleasesDGX agent

arXiv:2605.29089v1 Announce Type: cross Abstract: Recent reinforcement learning (RL) post-training approaches primarily optimize the final output policy using sparse outcome-level rewards, while large

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

Model ReleasesDGX agent

arXiv:2605.29833v1 Announce Type: new Abstract: As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdi

OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources

Model ReleasesDGX agent

arXiv:2605.29250v1 Announce Type: cross Abstract: Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graph

On-Policy Replay for Continual Supervised Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.29495v1 Announce Type: new Abstract: Continual supervised fine-tuning (SFT) is the de facto recipe for adapting large language models (LLMs) to a stream of downstream tasks, but it suffers

On the Construction and Implications of Low-Loss Valleys in LoRA-based Bayesian Inference

Model ReleasesDGX agent

arXiv:2605.29580v1 Announce Type: new Abstract: While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic

On the remarkable return on capital potential for Starlink on Starship. Including customer acquisition cost, ground station capex, and an ex…

Model ReleasesDGX agent

On the remarkable return on capital potential for Starlink on Starship. Including customer acquisition cost, ground station capex, and an expendable top stage, we think SpaceX should be able to launch

OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction

Model ReleasesDGX agent

arXiv:2605.30247v1 Announce Type: new Abstract: Drug synergy prediction (DSP) aims to identify efficacious drug combinations under various cellular contexts with different targets. However, the contin

OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories

Model ReleasesDGX agent

arXiv:2605.29253v1 Announce Type: new Abstract: Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambi

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

Model ReleasesDGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

Optimizing Latent Representations for Robust Building Damage Assessment Onboard Earth Observation Satellites

Model ReleasesDGX agent

arXiv:2605.29575v1 Announce Type: new Abstract: Rapid identification of damaged buildings after natural disasters or on war areas is crucial to support emergency response and prioritize interventions.

OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation

Model ReleasesDGX agent

arXiv:2605.29829v1 Announce Type: new Abstract: Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient par

Orthogonal Concept Erasure for Diffusion Models

Model ReleasesDGX agent

arXiv:2605.28902v1 Announce Type: new Abstract: Concept erasure has emerged as a promising approach to mitigate undesired or unsafe content in diffusion models, yet existing methods still face signifi

OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment

Model ReleasesDGX agent

arXiv:2605.29900v1 Announce Type: new Abstract: Contrastive learning is effective for aligning paired views or modalities, but alignment beyond two modalities remains non-trivial and comparatively und

Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies

Model ReleasesDGX agent

arXiv:2605.30148v1 Announce Type: cross Abstract: Evolution Strategies (ES) has recently emerged as a competitive alternative to reinforcement learning (RL) for large language model (LLM) fine-tuning,

Parallax: Parameterized Local Linear Attention for Language Modeling

Model ReleasesDGX agent

arXiv:2605.29157v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become the central paradigm in artificial intelligence, yet the core computational primitive of attention has remain

Parameter-Efficient Subspace Decoupling ViT for Mitigating Multi-Task Negative Transfer in Histological Scoring

Model ReleasesDGX agent

arXiv:2605.29852v1 Announce Type: new Abstract: Histological scoring is essential for diagnosing Non-Alcoholic Fatty Liver Disease (NAFLD), yet its automation remains challenging due to the high annot

ParaTool: Shifting Tool Representations from Context to Parameters

Model ReleasesDGX agent

arXiv:2605.29561v1 Announce Type: new Abstract: Tool calling extends large language models (LLMs) by enabling grounded interaction with external executable interfaces, thereby supporting environment-c

Parse PDFs in the browser, or the edge, in milliseconds Our LiteParse WASM package can be literally run anywhere, from cloudflare workers, m…

Model ReleasesDGX agent

Parse PDFs in the browser, or the edge, in milliseconds Our LiteParse WASM package can be literally run anywhere, from cloudflare workers, mobile runtimes, to the browser. Starter template for Cloudfl

Personalized Turn-Level User Conversation Satisfaction Benchmark

Model ReleasesDGX agent

arXiv:2605.29711v1 Announce Type: cross Abstract: User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user

PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

Model ReleasesDGX agent

arXiv:2605.29710v1 Announce Type: new Abstract: Real-world evaluation of vision-language-action (VLA) policies still rests on binary success rate at a fixed timeout with N le 25 rollouts per condition

PhoneWorld: Scaling Phone-Use Agent Environments

Model ReleasesDGX agent

arXiv:2605.29486v1 Announce Type: cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Ex

Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software

Model ReleasesDGX agent

arXiv:2605.30353v1 Announce Type: new Abstract: Are AI agents tools, co-authors, or researchers? We present a quantified case study (N=1): a physicist supervising an AI coding agent (Claude Code, Sonn

pibot is now running fully local, using parakeet for STT, qwen3-tts for TTS, and Qwen 3.6 as the local multi-modal LLM via llama.cpp. The ST…

Model ReleasesDGX agent

pibot is now running fully local, using parakeet for STT, qwen3-tts for TTS, and Qwen 3.6 as the local multi-modal LLM via llama.cpp. The STT and TTS inference engines are Rust/mlx-c based. Ported fro

Planning with the Views via Scene Self-Exploration

Model ReleasesDGX agent

arXiv:2605.29563v1 Announce Type: new Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1)understandin

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.29299v1 Announce Type: cross Abstract: Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cos

PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers

Model ReleasesDGX agent

arXiv:2605.30094v1 Announce Type: new Abstract: Poker is a landmark challenge for artificial intelligence. The dominant approach relies on equilibrium solvers built on counterfactual regret minimizati

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

Model ReleasesDGX agent

arXiv:2605.29815v1 Announce Type: new Abstract: The growing number of submitted papers has motivated the exploration of Large Language Models (LLMs) as a means to support and augment the peer review p

Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit

Model ReleasesDGX agent

arXiv:2605.28873v1 Announce Type: new Abstract: This is a planning-method note with an unpaired pilot audit. We adapt the classical paired-binary sample-size calculation (Miettinen, 1968) to quantizat

Predicting Causal Effects from Natural Language Queries using Structured Representations

Model ReleasesDGX agent

arXiv:2605.29631v1 Announce Type: cross Abstract: Randomized controlled trials are a cornerstone of medicine and the social sciences as they enable reliable estimates of causal effects. However, they

Prescribe-then-Select: Adaptive Policy Selection for Contextual Stochastic Optimization

Model ReleasesDGX agent

arXiv:2509.08194v2 Announce Type: replace Abstract: We address the problem of policy selection in contextual stochastic optimization (CSO), where covariates are available as contextual information and

Prioritize the Process, Not Just the Outcome: Rewarding Latent Thought Trajectories Improves Reasoning in Looped Language Models

Model ReleasesDGX agent

arXiv:2602.10520v3 Announce Type: replace Abstract: Looped Language Models (LoopLMs) perform multi-step latent reasoning prior to token generation and outperform conventional LLMs on reasoning benchma

← Previous
1…194195196197198…377
Next →