AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,106 results
11 Aug 2026

ChatGPT and Gemini both just passed 1 billion users

Model ReleasesDGX agent

For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing produc

ChronoState: Hidden Elapsed-Time Conditioning for Temporal-State Action Selection in Frozen-Backbone Language Models

Model ReleasesDGX agent

arXiv:2608.09124v1 Announce Type: new Abstract: Temporal decisions in language-model systems often depend on both symbolic task state and elapsed wall-clock time, such as cache expiration, job complet

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.09164v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms.

Circuit Fine-Tuning for Compute-Efficient Transformer Adaptation

Model ReleasesDGX agent

arXiv:2608.08336v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become the de facto standard for adapting Vision Transformers (ViTs) to downstream tasks. While parameter cou

CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

Model ReleasesDGX agent

arXiv:2608.09374v1 Announce Type: new Abstract: Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, sel

Claude will apply invisible watermarks to AI text and images

Model ReleasesDGX agent

Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. 'Generated text will carry embedded

Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building te…

Model ReleasesDGX agent

Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can b

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

Model ReleasesDGX agent

arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investiga

Closing the loop in learning with missing data

Model ReleasesDGX agent

arXiv:2608.09030v1 Announce Type: cross Abstract: What should a machine learning model learn when data is missing during training? We look at the learning process from a dynamical systems perspective,

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primari

CodecArena: Codec Quality Assessment via Visual Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.09139v1 Announce Type: new Abstract: Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly opt

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

Model ReleasesDGX agent

arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise

COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping

Model ReleasesDGX agent

arXiv:2608.07570v1 Announce Type: cross Abstract: Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-

Compiling and Benchmarking Task-State Horizons for Embodied Agents

Model ReleasesDGX agent

arXiv:2608.08036v1 Announce Type: new Abstract: Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

Model ReleasesDGX agent

arXiv:2608.07584v1 Announce Type: new Abstract: Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such t

Contamination Means Overestimation? A Fine-Grained Empirical Study in Code Intelligence

Model ReleasesDGX agent

arXiv:2506.02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training

Model ReleasesDGX agent

arXiv:2608.08224v1 Announce Type: new Abstract: Reinforcement learning post-training unlocks complex reasoning in LLMs. Yet benchmark scores reveal only whether a model improved, not what changed insi

Conversation as Measurement in Clinical Encounters: Observable Phase Structure, Partially Observable Patient State

Model ReleasesDGX agent

arXiv:2608.08868v1 Announce Type: new Abstract: Many modern AI systems analyze conversational traces to infer aspects of human interaction and state, implicitly assuming that such information is recov

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2608.08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the 'right' answer, but how they decide what matters most when moral principle

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.09324v1 Announce Type: new Abstract: On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding th

CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting

Model ReleasesDGX agent

arXiv:2608.07693v1 Announce Type: cross Abstract: Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short observation his

Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge

Model ReleasesDGX agent

arXiv:2608.07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures. Howe

Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

Model ReleasesDGX agent

arXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on

CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge

Model ReleasesDGX agent

arXiv:2604.03374v2 Announce Type: replace-cross Abstract: Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, and c

Cross-Model Humor Preference Modeling with Cards Against Humanity

Model ReleasesDGX agent

arXiv:2608.07481v1 Announce Type: cross Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Model ReleasesDGX agent

arXiv:2608.09766v1 Announce Type: cross Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evalu

Curriculum Generation under Structured Parametric Environments for Robust Navigation Policies

Model ReleasesDGX agent

arXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f

DarwinX: Evolving Agent Harnesses Through Natural Selection

Model ReleasesDGX agent

arXiv:2608.07545v1 Announce Type: cross Abstract: An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops alrea

Decentralized Nonconvex Composite Federated Learning with Gradient Tracking and Momentum

Model ReleasesDGX agent

arXiv:2504.12742v2 Announce Type: replace Abstract: Decentralized Federated Learning (DFL) enables collaborative model training without relying on a central server. When local objectives are nonconvex

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

Model ReleasesDGX agent

arXiv:2608.09900v1 Announce Type: new Abstract: Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably wa

Decoupled Descent: Enforcing Exact Train-Test Error Tracking Via AMP Onsager Corrections [R]

Model ReleasesDGX agent

Link: https://arxiv.org/pdf/2604.27883 Hi, Most of use are familiar with the headache of training a neural network using gradient descent where the training error may go to zero but the test error may

DeepSeek-V4-Flash-0731 (284B MoE) at 75 tok/s on 2× DGX Spark — full recipe, 11 gotchas, reboot-proof cluster, Codex CLI integration

Model ReleasesDGX agent

Spent two nights getting deepseek-ai/DeepSeek-V4-Flash-0731 (284B MoE, 13B active, native FP4/FP8, 1M context) running production-grade on two DGX Sparks connected by one QSFP DAC cable. Everything —

DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide

Model ReleasesDGX agent

Been benchmarking DSv4 Flash 0731 on a Flow Z13 (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151, 128GB LPDDR5X) for the past week. Figured I'd share what actually works and what doesn't — there are a lot o

DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO…

Model ReleasesDGX agent

DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO, then deploy the fine-tuned model on Together AI for produc

DeepSeek-V4-Flash acting as my Linux sysadmin

Model ReleasesDGX agent

I'm very happy with some Linux admin tasks I'm throwing at a locally running DeepSeek. My request was simple, check why 'samples' folder is taking more and more space on one of the machines on my LAN,

Deploying Anthropic Claude apps gateway for AWS for enterprise workloads

Model ReleasesDGX agent

Claude apps gateway is a self-hosted governance layer between Claude Code and Claude Desktop and Amazon Bedrock or Claude Platform on AWS. This post presents a production reference deployment covering

Depth-Aware Implicit Neural Representation Priors for 3D Gravity Inversion

Model ReleasesDGX agent

arXiv:2608.08959v1 Announce Type: new Abstract: Gravimetry images subsurface density contrasts associated with geological structures, geothermal systems, and intrusive bodies. Recovering a three-dimen

Describe-to-Score: A text-guided framework for image complexity assessment

Model ReleasesDGX agent

arXiv:2509.16609v2 Announce Type: replace Abstract: Accurately assessing image complexity (IC) is essential for many vision tasks, yet existing approaches rely almost exclusively on visual features an

Detecting Clear Contact Lenses for Iris Recognition: A Two-Stage Mask-Guided Attention Approach

Model ReleasesDGX agent

arXiv:2608.08977v1 Announce Type: cross Abstract: This work focuses on the impact and detection of clear contact lenses in the context of iris recognition. While the detection of cosmetic or patterned

DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?

Model ReleasesDGX agent

arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code

Diagnosing as Cardiologists Do: ECG Agents with Doctor-Grounded Priors for Clinical Reasoning Across Diseases and Populations

Model ReleasesDGX agent

arXiv:2608.09053v1 Announce Type: cross Abstract: Cardiologists interpret electrocardiograms by localizing waveform components, measuring rhythm and interval patterns, and translating these structured

Diffuse the object, keep its label: curating detector training data from a few unlabeled photographs via VLM-built 3D vegetation scenes

Model ReleasesDGX agent

arXiv:2608.09691v1 Announce Type: new Abstract: Labeled images of small objects hidden in vegetation are scarce, and detectors trained on them generalize poorly across sites. Rather than reusing label

Diminishing Returns of Intelligence: The Non-Linear Relationship Between LLM Scale and User Perception in Short-Duration Open-Ended Social Human-Robot Interactions

Model ReleasesDGX agent

arXiv:2608.08320v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to drive embodied social agents, yet it remains unclear whether larger models improve user perception

Directional-Clamp PPO

Model ReleasesDGX agent

arXiv:2511.02577v2 Announce Type: replace Abstract: Proximal Policy Optimization (PPO) is widely regarded as one of the most successful deep reinforcement learning algorithms, known for its robustness

Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity Optimization

Model ReleasesDGX agent

arXiv:2608.08523v1 Announce Type: new Abstract: Multimodal embodied agents are increasingly required to solve long-horizon tasks by integrating visual observations, textual goals, and interaction hist

Disentangling Co-Occurring Retinal Pathologies with Saliency-Guided Sparse Expert Routing

Model ReleasesDGX agent

arXiv:2608.09752v1 Announce Type: new Abstract: Retinal fundus images frequently exhibit multiple co-occurring pathologies, yet standard deep learning classifiers apply static, identical computation t

DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

Model ReleasesDGX agent

arXiv:2608.08878v1 Announce Type: cross Abstract: Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequen

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

Model ReleasesDGX agent

arXiv:2608.08029v1 Announce Type: cross Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-

DocAtlas: Long-Document Understanding as Mutable-State Interaction

Model ReleasesDGX agent

arXiv:2608.07527v1 Announce Type: cross Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-a

Does a Toehold Make a Bidder Bolder? Preemption and Multiplicity in Multi-Round Takeover Auctions

Model ReleasesDGX agent

arXiv:2608.08407v1 Announce Type: cross Abstract: A bidder can quietly buy a stake in a company before making an offer for it. That stake, a toehold, is supposed to pay for itself twice: it makes the

Domain-Aware Pruning: Sparsity and Domain Generalization via Regularized Probabilistic Masking

Model ReleasesDGX agent

arXiv:2608.08624v1 Announce Type: new Abstract: Domain generalization (DG) and neural network pruning are conventionally treated as distinct objectives, targeting out-of-distribution (OOD) robustness

Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization

Model ReleasesDGX agent

arXiv:2608.09043v1 Announce Type: cross Abstract: Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on its own. We

DualCert: A Solver for the Traveling Salesman Problem with Constraint-Coupled Learning

Model ReleasesDGX agent

arXiv:2608.09042v1 Announce Type: new Abstract: Large traveling salesman problem (TSP) instances require a solver to allocate limited computation while preserving the validity of its outputs. Existing

Effect of Abstractions and Prompting Strategies on LLM-Guided High-Performance Optimizations

Model ReleasesDGX agent

arXiv:2608.08085v1 Announce Type: cross Abstract: Code performance optimization is a vital aspect of modern software development, as it enables faster response times and reduced resource usage. These

Efficient Fine-Tuning of DINOv3 Pretrained on Natural Images for Atypical Mitotic Figure Classification

Model ReleasesDGX agent

arXiv:2508.21041v4 Announce Type: replace-cross Abstract: Atypical mitotic figures (AMFs) indicate abnormal cell division associated with poor prognosis. Their detection remains difficult due to low p

Efficient Human-Contact Representation for Human-Scene Interaction

Model ReleasesDGX agent

arXiv:2608.09388v1 Announce Type: new Abstract: Human-scene interaction is an active research topic with several industrial applications in virtual reality, gaming, robotics, and surveillance. Despite

Eikonal Regularisation in Physics-Informed Neural Networks for Three-Dimensional Level-Set Advection: Transferability of Two-Dimensional Design Principles

Model ReleasesDGX agent

arXiv:2608.08322v1 Announce Type: cross Abstract: Physics-informed neural networks applied to the level-set formulation of interface advection commonly augment the residual and initial-condition losse

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

Model ReleasesDGX agent

arXiv:2608.09548v1 Announce Type: cross Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that or

ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making

Model ReleasesDGX agent

arXiv:2608.09024v1 Announce Type: new Abstract: Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prior

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models

Model ReleasesDGX agent

arXiv:2608.09189v1 Announce Type: new Abstract: Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Model

← Previous
1…56789…369
Next →