AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,280 results
7 Aug 2026

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

Model ReleasesDGX agent

arXiv:2608.05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain beh

M^3R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

Model ReleasesDGX agent

arXiv:2608.05817v1 Announce Type: new Abstract: Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visu

MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.16284v2 Announce Type: replace Abstract: Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communicat

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

Model ReleasesDGX agent

arXiv:2608.05850v1 Announce Type: cross Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, it

MASS: Multiplayer World Models with Authoritative Shared State

Model ReleasesDGX agent

arXiv:2608.06257v1 Announce Type: new Abstract: Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redunda

Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents

Model ReleasesDGX agent

arXiv:2606.21140v2 Announce Type: replace-cross Abstract: Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs d

Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers

Model ReleasesDGX agent

arXiv:2608.05472v1 Announce Type: cross Abstract: Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping

MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers

Model ReleasesDGX agent

arXiv:2608.05878v1 Announce Type: new Abstract: Text-to-image diffusion transformers learn about objects and scenes by learning to generate them, making them strong candidates for training-free zero-s

MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

Model ReleasesDGX agent

arXiv:2608.06253v1 Announce Type: new Abstract: Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed Me

Mind the Gaps: Mixture-of-Minds for Human Simulation

Model ReleasesDGX agent

arXiv:2608.06115v1 Announce Type: new Abstract: Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the l

MoCA: Implicit Social Context Analysis

Model ReleasesDGX agent

arXiv:2608.05825v1 Announce Type: new Abstract: Human social communication, such as affection and intent, is often conveyed in highly implicit ways, where underlying meanings are expressed through ind

Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra)

Model ReleasesDGX agent

Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5, where I had Claude Fable 5 build a full working game

MS-MLB: An Open Machine Learning Benchmark for Blood-Based MS Classification

Model ReleasesDGX agent

arXiv:2608.05196v1 Announce Type: new Abstract: Multiple sclerosis (MS) is diagnosed through clinical assessment, magnetic resonance imaging, laboratory evidence when appropriate, and exclusion of bet

Multi-Representation Geometric Hierarchy Fusion: An Implicit-Submap Driven Framework for Resilient 3D Place Recognition

Model ReleasesDGX agent

arXiv:2506.14243v4 Announce Type: replace Abstract: LiDAR-based place recognition is critical for long-term autonomous driving without GPS. Existing handcrafted feature methods face dual limitations.

My issue with Artificial Analysis's 'intelligence index'

Model ReleasesDGX agent

I swear AA is not the bipartisan they so claim. An open source mode (Qwen 3.8 max) was number 1 on the agentic index, then they just so happen to launch 'v4.1.1' of their index in which they just adju

NavTrust: Benchmarking Trustworthiness for Embodied Navigation

Model ReleasesDGX agent

arXiv:2603.19229v2 Announce Type: replace-cross Abstract: There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural language i

NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering

Model ReleasesDGX agent

arXiv:2608.06292v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. H

Neuro-Symbolic Closed-Loop Control of Laser Powder Bed Fusion with an In-Loop Ontology

Model ReleasesDGX agent

arXiv:2608.05773v1 Announce Type: new Abstract: A geometry-conditioned, neuro-symbolic closed-loop architecture is proposed for laser powder bed fusion, in which a standards-aligned ontology operates

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

Model ReleasesDGX agent

arXiv:2608.05539v1 Announce Type: new Abstract: Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D obj

One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs

Model ReleasesDGX agent

arXiv:2512.14751v3 Announce Type: replace-cross Abstract: Finetuning pretrained large language models (LLMs) has become the standard paradigm for developing downstream applications. However, its secur

OpenAI puts the brakes on a new model because it’s supposedly too powerful

Model ReleasesDGX agent

OpenAI says it is pausing 'internal activities' around an in-development AI model, Astra, because it doesn't yet meet new security standards the company is putting in place. The announcement follows i

Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational Hardening

Model ReleasesDGX agent

arXiv:2608.05944v1 Announce Type: cross Abstract: We report operational experience full-fine-tuning a 32.76B-parameter dense model (Qwen3-32B) on 16 x NVIDIA B300 (two nodes, FSDP / ZeRO-3) -- among t

OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality

Model ReleasesDGX agent

arXiv:2608.05263v1 Announce Type: new Abstract: Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipelin

Otter: A Time-Aware, History-Conditioned Human Chess AI

Model ReleasesDGX agent

arXiv:2608.05206v1 Announce Type: new Abstract: Otter is a 15.3M-parameter human chess AI that predicts human move selection by modeling play as a time-aware, sequential process rather than treating e

Outstanding cost-to-performance from DeepSeek GPT-5.6 Luna (Max) performance for a 1/4th of the cost on ARC-AGI

Model ReleasesDGX agent

Outstanding cost-to-performance from DeepSeek GPT-5.6 Luna (Max) performance for a 1/4th of the cost on ARC-AGI DeepSeek V4 Flash from @deepseek_ai on ARC-AGI (Verified): - ARC-AGI-2: 61.4%, 0.04/task

Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection

Model ReleasesDGX agent

arXiv:2604.04444v2 Announce Type: replace Abstract: Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-trainin

Perfect reconstruction of sparse signals using nonconvexity control and one-step RSB message passing

Model ReleasesDGX agent

arXiv:2512.17426v2 Announce Type: replace-cross Abstract: We consider sparse signal reconstruction via minimization of the smoothly clipped absolute deviation (SCAD) penalty, and develop one-step repl

PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs

Model ReleasesDGX agent

arXiv:2601.21124v2 Announce Type: replace-cross Abstract: Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audi

Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations

Model ReleasesDGX agent

arXiv:2604.17359v2 Announce Type: replace-cross Abstract: Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no real

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

Model ReleasesDGX agent

arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden

Position: It's Time to Optimize LLMs for Self-Consistency

Model ReleasesDGX agent

arXiv:2608.05188v1 Announce Type: cross Abstract: Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition

Project2Task: Graph-Guided Project-Level Planning for Autonomous Research

Model ReleasesDGX agent

arXiv:2608.05225v1 Announce Type: new Abstract: Research agents can increasingly search literature, propose hypotheses, generate code, run experiments, and draft manuscripts from a single topic. Howev

Qwen 3.6 27B flags/settings in llama.cpp

Model ReleasesDGX agent

I run the following on a 5090 and have been okay with its performance, it does most things somewhere 80-100 t/s, though that can slow down at full 262k context - more like 40 t/s at times. I use it pr

Rectifying Geometric Misalignment: Online Source-Free Adaptation for Class-Imbalanced EEG

Model ReleasesDGX agent

arXiv:2608.05315v1 Announce Type: new Abstract: Electroencephalography (EEG) based Brain-Computer Interfaces (BCIs) often require unsupervised domain adaptation (UDA) to generalize across subjects and

Recursive Synthesis for Long-Horizon Terminal Tasks

Model ReleasesDGX agent

arXiv:2608.05466v1 Announce Type: new Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because ea

Retailers are updating their websites to rank highly in chatbot results, while making sure purchases are done on their own sites to collect customer data (Arriana McLymore/Reuters)

Model ReleasesDGX agent

Arriana McLymore / Reuters: Retailers are updating their websites to rank highly in chatbot results, while making sure purchases are done on their own sites to collect customer data — As shoppers incr

Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation

Model ReleasesDGX agent

arXiv:2608.06221v1 Announce Type: cross Abstract: Learning from demonstration (LfD) provides a developmental framework through which robots can develop motor skills by observing and imitating human dy

Robust Native Language Identification through Agentic Decomposition

Model ReleasesDGX agent

arXiv:2509.16666v2 Announce Type: replace Abstract: Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual

Runtime Observability for Heterogeneous Attention Memory

Model ReleasesDGX agent

arXiv:2608.05863v1 Announce Type: new Abstract: Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different

SafeDivertor: Faithful Divertor Heat Flux Reconstruction from Macroscopic Plasma State Signals via Time-Frequency Prior Exploitation

Model ReleasesDGX agent

arXiv:2608.05669v1 Announce Type: cross Abstract: Divertor heat-flux analysis is essential for understanding plasma-wall interactions and protecting plasma-facing components in magnetic-confinement fu

Scaffold-Mediated Post-Training: Co-Evolving Model Parameters and Procedural Scaffold Graphs

Model ReleasesDGX agent

arXiv:2608.05156v1 Announce Type: new Abstract: Post-training of large language models optimizes only parameters, while inference-time procedural scaffolds are typically designed independently of para

Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

Model ReleasesDGX agent

arXiv:2608.06167v1 Announce Type: new Abstract: We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by aut

Scientific Machine Learning of Chaotic Systems Learns Reduced-Order Equations for Neural Populations

Model ReleasesDGX agent

arXiv:2507.03631v4 Announce Type: replace Abstract: Extracting interpretable mathematical models from complex dynamical systems is difficult, especially for chaotic dynamics observed with noisy experi

SEAM: Global consistency beyond local accuracy in scientific machine learning

Model ReleasesDGX agent

arXiv:2608.05702v1 Announce Type: new Abstract: Scientific machine learning commonly validates models at the level of a subdomain, a benchmark split, or an explanation for one prediction. Yet such loc

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

Model ReleasesDGX agent

arXiv:2608.05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning error

Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?

Model ReleasesDGX agent

arXiv:2608.05864v1 Announce Type: new Abstract: Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are l

SemiAdapt-Instruct: Extensible Instruction Tuning via Latent Domain-Specialised Adapters

Model ReleasesDGX agent

arXiv:2608.05161v1 Announce Type: new Abstract: Instruction-tuned LLMs are deployed into environments where domains evolve, yet extending a fine-tuned model's capabilities without full retraining rema

Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?

Model ReleasesDGX agent

Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit

Shrinking the Generation-Verification Gap with Weak Verifiers

Model ReleasesDGX agent

arXiv:2506.18203v3 Announce Type: replace Abstract: Verifiers can improve language model capabilities by scoring and ranking responses from generated candidates. Currently, high-quality verifiers are

Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support

Model ReleasesDGX agent

arXiv:2608.05151v1 Announce Type: cross Abstract: Wastewater operators need answers grounded in how their plant's variables interact and how fast effects propagate, not in generic pretraining text, wh

SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation

Model ReleasesDGX agent

arXiv:2608.05628v1 Announce Type: new Abstract: Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-w

SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation

Model ReleasesDGX agent

arXiv:2608.05970v1 Announce Type: cross Abstract: Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on roboti

SkillTFM: Gated Skill Evolution for Training-Free Adaptation of Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2608.06137v1 Announce Type: new Abstract: Tabular data are ubiquitous in real-world applications and are crucial for data-driven prediction and decision-making across science, industry, finance,

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

Model ReleasesDGX agent

arXiv:2608.05573v1 Announce Type: new Abstract: LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to veri

StepReflect: Structured UI Transition Reflection for Mobile GUI Agents

Model ReleasesDGX agent

arXiv:2608.05587v1 Announce Type: new Abstract: Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal r

Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data

Model ReleasesDGX agent

arXiv:2603.00059v3 Announce Type: replace-cross Abstract: How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

Model ReleasesDGX agent

arXiv:2608.05703v1 Announce Type: new Abstract: Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-s

Subliminal Learning is Non-Semantic Distillation

Model ReleasesDGX agent

arXiv:2608.05734v1 Announce Type: new Abstract: Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or behavior from a

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

Model ReleasesDGX agent

arXiv:2608.05785v1 Announce Type: cross Abstract: Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fund

TAU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding

Model ReleasesDGX agent

arXiv:2608.05699v1 Announce Type: new Abstract: Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unf

← Previous
1…2223242526…372
Next →