AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,323 results
31 Jul 2026

LLMs struggle to simulate human belief updates in controlled environments

Model ReleasesDGX agent

arXiv:2607.28347v1 Announce Type: new Abstract: LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been

LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning

Model ReleasesDGX agent

arXiv:2607.28135v1 Announce Type: new Abstract: Machine learning for combinatorial optimization typically relies on neural constructors trained via reinforcement learning on large offline datasets for

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.27806v1 Announce Type: new Abstract: In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal in

Looped Transformers with Source-Centered State Evolution

Model ReleasesDGX agent

arXiv:2607.27656v1 Announce Type: cross Abstract: Looped Transformers create a useful train- and test-time compute axis by reusing the same Transformer block over recurrent depth, increasing effective

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts

Model ReleasesDGX agent

arXiv:2607.27787v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on 'cliff' prompts-those on which

MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking

Model ReleasesDGX agent

arXiv:2607.17751v2 Announce Type: cross Abstract: We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, desi

MatCreatioNN: Machine learning-guided computational discovery of photocatalysts for environmental applications

Model ReleasesDGX agent

arXiv:2607.27295v1 Announce Type: cross Abstract: The rational design of photocatalysts for environmental remediation and CO2 conversion remains limited by the high computational cost and sparse exper

Measuring Alignment With Reader Highlights Net of Position and Length

Model ReleasesDGX agent

arXiv:2607.27739v1 Announce Type: cross Abstract: Context compression discards most of a document before a language model reads it, and is normally evaluated by downstream task accuracy - which makes

MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models

Model ReleasesDGX agent

arXiv:2502.20780v2 Announce Type: replace-cross Abstract: The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which t

Meituan just dropped LongCat-Flash-Lite-Sparse

Model ReleasesDGX agent

It’s an MoE with ~3B active params and a 30B n-gram lookup table offloaded to RAM for fast 256k context on a 24GB GPU. Reminds me of Gemma 4’s PLE trick. Initial analysis suggest it wont be replacing

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Model ReleasesDGX agent

arXiv:2607.27919v1 Announce Type: new Abstract: Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independent

MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair

Model ReleasesDGX agent

arXiv:2607.27080v1 Announce Type: cross Abstract: Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instr

MiniMax H3 discussion

Model ReleasesDGX agent

https://x.com/MiniMax_AI/status/2083008095488516262 (says it's coming in a few days) it can do text-to-image, and image editing everyone seems to be mostly hyped about the video generation part (I am

MiniMax H3: Open-weight multimodel video model

Model ReleasesDGX agent

Just saw this posted by Fal.ai and then by Hailuo themselves, the next video model will be open weight released! Here's the blurb and link to to the feature post: Today, we're launching MiniMax H3, a

Minimax-H3 video model released, open weights coming in the next few days

Model ReleasesDGX agent

https://x.com/MiniMax_AI/status/2083006198828417501?s=20 Quote from their article: Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context acro

Minimum VRAM GPU to run DeepSeek-V4-Flash-0731 Q4_K_XL at around 30 t/s ?

Model ReleasesDGX agent

Hello guys, I'm curious about running DeepSeek-V4-Flash-0731 locally. Since it’s a Mixture of Experts (MoE) model with only 13B active parameters, I was hoping the VRAM requirements might be manageabl

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

Model ReleasesDGX agent

arXiv:2607.27109v2 Announce Type: cross Abstract: With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-gra

MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

Model ReleasesDGX agent

arXiv:2607.27895v1 Announce Type: cross Abstract: Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological s

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.27637v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or sh

Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models

Model ReleasesDGX agent

arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We p

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

Model ReleasesDGX agent

arXiv:2511.12449v3 Announce Type: replace Abstract: Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challen

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

Model ReleasesDGX agent

arXiv:2607.28274v1 Announce Type: new Abstract: Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicat

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Model ReleasesDGX agent

arXiv:2607.27616v1 Announce Type: new Abstract: Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into share

MSCM-net: A hyperspectral image classiffcation method based on multi-scale convolution and Mamba

Model ReleasesDGX agent

arXiv:2607.28277v1 Announce Type: new Abstract: Hyperspectral imaging is widely used in remote sensing and engineering. Therefore, research on its classification methods is crucial. While CNN and Tran

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Model ReleasesDGX agent

arXiv:2607.26465v1 Announce Type: new Abstract: Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform seq

Neat work on long-horizon agents. Splitting a hard task across agents is typically how standard multi-agent work. The usual design lets them…

Model ReleasesDGX agent

Neat work on long-horizon agents. Splitting a hard task across agents is typically how standard multi-agent work. The usual design lets them exchange findings only at phase boundaries, through staged

Neural Network-Assisted CLEAN for Channel Modeling in Low-SNR Regimes

Model ReleasesDGX agent

arXiv:2607.27450v1 Announce Type: new Abstract: Accurate multipath parameter estimation is critical for modern wireless communication systems, particularly in challenging low-SNR environments. Traditi

New research from Microsoft. This one is on training computer-use agents at scale. Recent pipelines generate synthetic environments in bulk,…

Model ReleasesDGX agent

New research from Microsoft. This one is on training computer-use agents at scale. Recent pipelines generate synthetic environments in bulk, which moved the bottleneck from how many exist to what is i

Now Suddenly too many choices for DGX Spark with Qwen 3.5 122B . What would be the next upgrade?

Model ReleasesDGX agent

Laguna 2.1 at NVFP4 Deepseek v4 at Q2 Inkling-Small at IQ3 Which models you guys running now ? How it compares to 122b? Upcoming in few days : Ling 3.0 124B (Could be new king) LongCat 69B A3B ( very

Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA

Model ReleasesDGX agent

arXiv:2607.27566v1 Announce Type: new Abstract: Multi-frame medical VQA appears to reward increasingly complex adaptation: controller-style inference, localization-aware reranking, static hard-negativ

OK, GPT-5.6 Luna is a bit of a beast. Given the 80% price drop today I decided to try it in Datasette Agent, and it's furiously quick and ge…

Model ReleasesDGX agent

OK, GPT-5.6 Luna is a bit of a beast. Given the 80% price drop today I decided to try it in Datasette Agent, and it's furiously quick and generates all the SQL, HTML and JavaScript (for Datasette Apps

On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems

Model ReleasesDGX agent

arXiv:2607.28080v1 Announce Type: cross Abstract: We extend a recently introduced Entropy-Optimal Manifold Clustering (EOMC) to allow for a joint simultaneous identification of subsets and subspaces o

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

Model ReleasesDGX agent

arXiv:2607.27902v1 Announce Type: new Abstract: Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as

Open Source Ternary LLM Engine in Rust/CUDA for Quantization, Serving, and Training of models on consumer GPUs, called Tritium (Apache 2.0)

Model ReleasesDGX agent

This post was not written by a clanker. Hey guys, I'm a comp sci major who wanted to introduce a cool project I built for quantizing models to ternary (1.58 bit) with as minimal of loss as possible, a

OpenAI’s entire growth loop: new models, price cuts, and Tibo’s token resets

Model ReleasesDGX agent

OpenAI’s entire growth loop: new models, price cuts, and Tibo’s token resets major price cuts today: *80% drop for GPT-5.6 Luna, now 0.20 per million input tokens and 1.20 per million output *20% drop

Optimal Realistic Local AI for Most

Model ReleasesDGX agent

So you’ve got a 3090 or maybe even a 5090? Or more likely a 4060 8GB Ti. You wanna try local AI, you don’t know what it can/can’t do. 1) Install the best model you can. If you have a 3090 or a 5090, t

Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory

Model ReleasesDGX agent

arXiv:2607.28437v1 Announce Type: new Abstract: Molecular optimization is commonly performed under a limited oracle budget, which makes deciding what to evaluate as important as deciding what to gener

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Model ReleasesDGX agent

arXiv:2607.28545v1 Announce Type: new Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics,

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Model ReleasesDGX agent

arXiv:2607.28609v1 Announce Type: cross Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Veri

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

Model ReleasesDGX agent

arXiv:2607.27278v1 Announce Type: new Abstract: Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchm

Oxide and Friends: The Open Weight Revolution with Simon Willison

Model ReleasesDGX agent

Oxide and Friends: The Open Weight Revolution with Simon Willison On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild week we've had - with Kimi K3 show

PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball

Model ReleasesDGX agent

arXiv:2607.28623v1 Announce Type: new Abstract: We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body hum

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

Model ReleasesDGX agent

arXiv:2607.27378v1 Announce Type: new Abstract: Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-g

PCAP-LM: An LLM-Native Text Representation for TLS Bulk Traffic Analysis

Model ReleasesDGX agent

arXiv:2607.28100v1 Announce Type: cross Abstract: Large language models (LLMs) offer powerful reasoning capabilities for network traffic analysis, but standard capture formats and their textual equiva

Period.

Model ReleasesDGX agent

Period. If LoRA is underperforming, don't reach for more expensive full parameter fine-tuning right away. We ran three cheap tests (data coverage, optimization, rank) to see if we could close the gap

PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform's Perspective

Model ReleasesDGX agent

arXiv:2607.27265v1 Announce Type: new Abstract: Real-time bidding is central to computational advertising, comprising three elements: Supply Side Platform (SSP) selling ad impressions, Demand Side Pla

Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

Model ReleasesDGX agent

arXiv:2607.27764v1 Announce Type: new Abstract: Publishing private face recognition~(FR) training datasets is privacy-sensitive because faces expose identity information. Private FR training dataset p

Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry

Model ReleasesDGX agent

arXiv:2607.27224v1 Announce Type: cross Abstract: External and synthetic control arms (ECAs) are entering psychiatric drug development, but the field lacks a benchmark that evaluates the properties re

QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction

Model ReleasesDGX agent

arXiv:2607.28422v1 Announce Type: new Abstract: Fault-tolerant quantum computing (FTQC) relies on quantum error correction to suppress physical errors and preserve logical information at scale. In pra

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Model ReleasesDGX agent

arXiv:2607.28227v1 Announce Type: cross Abstract: GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision a

Recursive transformers for semiconductor thermo-mechanical reliability

Model ReleasesDGX agent

arXiv:2607.27251v1 Announce Type: new Abstract: Transformer-based surrogate models are increasingly used to replace expensive first-principles simulation in engineering design. But conventional transf

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy

Model ReleasesDGX agent

arXiv:2607.27782v1 Announce Type: new Abstract: Flow-matching Vision-Language-Action (VLA) policies have shown strong potential for robotic manipulation but often suffer from compounding errors caused

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Model ReleasesDGX agent

arXiv:2607.28509v1 Announce Type: new Abstract: Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference

Region-adaptable retrieval of coastal biogeochemical parameters from near-surface hyperspectral remote sensing reflectance using physics-aware meta-learning

Model ReleasesDGX agent

arXiv:2605.05623v2 Announce Type: replace Abstract: Hyperspectral in situ sensing has shown promise in retrieving aquatic biogeochemical (BGC) parameters, such as total suspended solids, dissolved org

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

Model ReleasesDGX agent

arXiv:2607.28008v1 Announce Type: new Abstract: Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthe

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

Model ReleasesDGX agent

arXiv:2607.28128v1 Announce Type: new Abstract: LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit

Revisiting Predictive Process Monitoring in the Age of Foundation Models: A Comparative Study of Sequence, Tabular, and LLM Approaches

Model ReleasesDGX agent

arXiv:2607.27797v1 Announce Type: new Abstract: Predictive process monitoring (PPM) leverages event logs to forecast the future of running process instances, for instance, predicting the next activity

RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design

Model ReleasesDGX agent

arXiv:2603.01229v3 Announce Type: replace Abstract: Robotic manipulation policies have made rapid progress in recent years, yet most existing approaches give limited consideration to memory capabiliti

Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

Model ReleasesDGX agent

arXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an

Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models

Model ReleasesDGX agent

arXiv:2607.27384v1 Announce Type: new Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure

← Previous
1…4748495051…373
Next →