AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

85,202Total entries
1Added by human
85,201Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,805 results
Model Releases

MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models

DGX agent

arXiv:2507.09574v3 Announce Type: replace-cross Abstract: Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requ

model-releasesarxiv-cs-ai
29 May 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Mind Your Tone: Does Tone Alter LLM Performance?

DGX agent

arXiv:2605.29027v1 Announce Type: new Abstract: The use of Large Language Models (LLMs) is proliferating, yet their performance is observed to vary based on prompting styles and tones. In this study,

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

DGX agent

arXiv:2605.29737v1 Announce Type: cross Abstract: LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code t

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

DGX agent

arXiv:2605.29360v1 Announce Type: new Abstract: Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that t

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Mitigating Hallucination in Vision-Language Models through Barrier-Regulated Adaptive Closed-form Steering

DGX agent

arXiv:2605.29881v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) often hallucinate objects that are not present in the input image, largely because visual grounding weakens as de

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions

DGX agent

arXiv:2605.29862v1 Announce Type: cross Abstract: AI-driven respiratory sound classification (RSC) is promising for automated pulmonary disease detection, yet multi-site deployment is hindered by inte

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

models underestimate how much work it takes (token usage) to accomplish a task, just like us

DGX agent

models underestimate how much work it takes (token usage) to accomplish a task, just like us 🧵 Claude-Opus-4.8 takes you too much tokens - but is this issue general across agents? Do agents know how m

model-releasesyohei-nakajima--x
29 May 2026
Model Releases

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing

DGX agent

arXiv:2605.22100v2 Announce Type: replace Abstract: Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information sys

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

DGX agent

arXiv:2605.29738v1 Announce Type: cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models

DGX agent

arXiv:2602.14399v2 Announce Type: replace Abstract: Multi-turn jailbreak attacks have proven effective against text-only large language models (LLMs), where malicious content is gradually introduced t

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Multimodal LLMs See Sentiment

DGX agent

arXiv:2508.16873v3 Announce Type: replace Abstract: Understanding how visual content conveys sentiment is increasingly important in a digital landscape dominated by imagery. However, sentiment percept

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization

DGX agent

arXiv:2605.29951v1 Announce Type: new Abstract: Understanding how harm emerges from interaction between otherwise benign image-text pairs requires intent-aware cross-modal reasoning beyond surface-lev

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

DGX agent

arXiv:2605.29300v1 Announce Type: cross Abstract: Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses ar

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs

DGX agent

arXiv:2605.29716v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising non-autoregressive generative paradigm. Given the prohibitive computational cost of

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

DGX agent

arXiv:2605.29685v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

No More K-means:Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval

DGX agent

arXiv:2605.30120v1 Announce Type: cross Abstract: Multi-vector retrieval (MVR) models, exemplified by ColBERT, have established new benchmarks in retrieval accuracy by preserving fine-grained token-le

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

DGX agent

arXiv:2605.29676v1 Announce Type: new Abstract: Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default languag

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Nvidia up 0.7%, on news that tokenmaxxing is dead and H200 rental prices are down. What an absurd time to be alive.

DGX agent

Nvidia up 0.7%, on news that tokenmaxxing is dead and H200 rental prices are down. What an absurd time to be alive. In the last 30 days alone: – Microsoft cancelled most of its Claude Code licenses, c

model-releasesgary-marcus--x
29 May 2026
Model Releases

OISD: On-Policy Internal Self-Distillation of Language Models

DGX agent

arXiv:2605.29089v1 Announce Type: cross Abstract: Recent reinforcement learning (RL) post-training approaches primarily optimize the final output policy using sparse outcome-level rewards, while large

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

DGX agent

arXiv:2605.29833v1 Announce Type: new Abstract: As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources

DGX agent

arXiv:2605.29250v1 Announce Type: cross Abstract: Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graph

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

On-Policy Replay for Continual Supervised Fine-Tuning

DGX agent

arXiv:2605.29495v1 Announce Type: new Abstract: Continual supervised fine-tuning (SFT) is the de facto recipe for adapting large language models (LLMs) to a stream of downstream tasks, but it suffers

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

On the Construction and Implications of Low-Loss Valleys in LoRA-based Bayesian Inference

DGX agent

arXiv:2605.29580v1 Announce Type: new Abstract: While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

On the remarkable return on capital potential for Starlink on Starship. Including customer acquisition cost, ground station capex, and an ex…

DGX agent

On the remarkable return on capital potential for Starlink on Starship. Including customer acquisition cost, ground station capex, and an expendable top stage, we think SpaceX should be able to launch

model-releaseselon-musk--x
29 May 2026
Model Releases

OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction

DGX agent

arXiv:2605.30247v1 Announce Type: new Abstract: Drug synergy prediction (DSP) aims to identify efficacious drug combinations under various cellular contexts with different targets. However, the contin

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories

DGX agent

arXiv:2605.29253v1 Announce Type: new Abstract: Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

DGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Optimizing Latent Representations for Robust Building Damage Assessment Onboard Earth Observation Satellites

DGX agent

arXiv:2605.29575v1 Announce Type: new Abstract: Rapid identification of damaged buildings after natural disasters or on war areas is crucial to support emergency response and prioritize interventions.

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation

DGX agent

arXiv:2605.29829v1 Announce Type: new Abstract: Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient par

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Orthogonal Concept Erasure for Diffusion Models

DGX agent

arXiv:2605.28902v1 Announce Type: new Abstract: Concept erasure has emerged as a promising approach to mitigate undesired or unsafe content in diffusion models, yet existing methods still face signifi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment

DGX agent

arXiv:2605.29900v1 Announce Type: new Abstract: Contrastive learning is effective for aligning paired views or modalities, but alignment beyond two modalities remains non-trivial and comparatively und

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies

DGX agent

arXiv:2605.30148v1 Announce Type: cross Abstract: Evolution Strategies (ES) has recently emerged as a competitive alternative to reinforcement learning (RL) for large language model (LLM) fine-tuning,

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Parallax: Parameterized Local Linear Attention for Language Modeling

DGX agent

arXiv:2605.29157v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become the central paradigm in artificial intelligence, yet the core computational primitive of attention has remain

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Parameter-Efficient Subspace Decoupling ViT for Mitigating Multi-Task Negative Transfer in Histological Scoring

DGX agent

arXiv:2605.29852v1 Announce Type: new Abstract: Histological scoring is essential for diagnosing Non-Alcoholic Fatty Liver Disease (NAFLD), yet its automation remains challenging due to the high annot

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

ParaTool: Shifting Tool Representations from Context to Parameters

DGX agent

arXiv:2605.29561v1 Announce Type: new Abstract: Tool calling extends large language models (LLMs) by enabling grounded interaction with external executable interfaces, thereby supporting environment-c

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Parse PDFs in the browser, or the edge, in milliseconds Our LiteParse WASM package can be literally run anywhere, from cloudflare workers, m…

DGX agent

Parse PDFs in the browser, or the edge, in milliseconds Our LiteParse WASM package can be literally run anywhere, from cloudflare workers, mobile runtimes, to the browser. Starter template for Cloudfl

model-releasesjerry-liu--x
29 May 2026
Model Releases

Personalized Turn-Level User Conversation Satisfaction Benchmark

DGX agent

arXiv:2605.29711v1 Announce Type: cross Abstract: User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

DGX agent

arXiv:2605.29710v1 Announce Type: new Abstract: Real-world evaluation of vision-language-action (VLA) policies still rests on binary success rate at a fixed timeout with N le 25 rollouts per condition

model-releasesarxiv-cs-ro
29 May 2026
Model Releases

PhoneWorld: Scaling Phone-Use Agent Environments

DGX agent

arXiv:2605.29486v1 Announce Type: cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Ex

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software

DGX agent

arXiv:2605.30353v1 Announce Type: new Abstract: Are AI agents tools, co-authors, or researchers? We present a quantified case study (N=1): a physicist supervising an AI coding agent (Claude Code, Sonn

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

pibot is now running fully local, using parakeet for STT, qwen3-tts for TTS, and Qwen 3.6 as the local multi-modal LLM via llama.cpp. The ST…

DGX agent

pibot is now running fully local, using parakeet for STT, qwen3-tts for TTS, and Qwen 3.6 as the local multi-modal LLM via llama.cpp. The STT and TTS inference engines are Rust/mlx-c based. Ported fro

model-releasesclem-delangue--x
29 May 2026
Model Releases

Planning with the Views via Scene Self-Exploration

DGX agent

arXiv:2605.29563v1 Announce Type: new Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1)understandin

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

DGX agent

arXiv:2605.29299v1 Announce Type: cross Abstract: Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cos

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers

DGX agent

arXiv:2605.30094v1 Announce Type: new Abstract: Poker is a landmark challenge for artificial intelligence. The dominant approach relies on equilibrium solvers built on counterfactual regret minimizati

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

DGX agent

arXiv:2605.29815v1 Announce Type: new Abstract: The growing number of submitted papers has motivated the exploration of Large Language Models (LLMs) as a means to support and augment the peer review p

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit

DGX agent

arXiv:2605.28873v1 Announce Type: new Abstract: This is a planning-method note with an unpaired pilot audit. We adapt the classical paired-binary sample-size calculation (Miettinen, 1968) to quantizat

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Predicting Causal Effects from Natural Language Queries using Structured Representations

DGX agent

arXiv:2605.29631v1 Announce Type: cross Abstract: Randomized controlled trials are a cornerstone of medicine and the social sciences as they enable reliable estimates of causal effects. However, they

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Prescribe-then-Select: Adaptive Policy Selection for Contextual Stochastic Optimization

DGX agent

arXiv:2509.08194v2 Announce Type: replace Abstract: We address the problem of policy selection in contextual stochastic optimization (CSO), where covariates are available as contextual information and

model-releasesarxiv-cs-lg
29 May 2026
← Previous
1…247248249250251…476
Next →