AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,620 results
3 Jun 2026

CREward: A Type-Specific Creativity Reward Model

Model ReleasesDGX agent

arXiv:2511.19995v2 Announce Type: replace Abstract: Creativity is a complex phenomenon. When it comes to representing and assessing creativity, treating it as a single undifferentiated quantity would

Critical evaluation of PINN for FWD inverse analysis and differentiable FEM as an alternative

Model ReleasesDGX agent

arXiv:2606.03210v1 Announce Type: cross Abstract: Automatic-differentiation-based inverse analysis methods, including physics-informed neural networks (PINNs) and differentiable programming, have rece

Cross-Lingual Token Arbitrage: Optimizing Code Agent Context Windows via Local LLM Preprocessing

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.03618v1 Announce Type: new Abstract: AI-assisted coding agents are bottlenecked by input-token cost. Two pathologies of raw human input drive much of this overhead: tokenization inefficienc

Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications

Model ReleasesDGX agent

arXiv:2603.01576v3 Announce Type: replace Abstract: Geo-Foundation Models (GFMs) have been evaluated across diverse Earth observation task including multiple domains and have demonstrated strong poten

DECA: Decentralizing Block-Wise Adam for Efficient LLM Full-Parameter Fine-Tuning on Non-IID Data

Model ReleasesDGX agent

arXiv:2606.03209v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) in privacy-sensitive and resource-constrained environments remains challenging. Since training data are often d

Decoupled Smart Contract Audits: Lightweight LLM Framework via Distillation and Aggregation

Model ReleasesDGX agent

arXiv:2606.03128v1 Announce Type: cross Abstract: Smart contracts face critical security challenges that require thorough auditing in decentralized web services. While Large Language Models (LLMs) hav

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials

Model ReleasesDGX agent

arXiv:2606.03951v1 Announce Type: new Abstract: Human experience in digital environments offers a vast, underexplored resource of authentic, untrimmed interactions that contain rich procedural knowled

DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration

Model ReleasesDGX agent

arXiv:2606.03103v1 Announce Type: new Abstract: Real-world professional desktop workflows in specialized creative and engineering software unfold over long horizons and often require human-in-the-loop

Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition

Model ReleasesDGX agent

arXiv:2606.03657v1 Announce Type: new Abstract: Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a functi

Diagnosis of Human Object Interaction Detectors for Real World Educational Applications

Model ReleasesDGX agent

arXiv:2606.02789v1 Announce Type: new Abstract: Human-object interaction (HOI) recognition is critical for automatically analyzing student behavior in complex educational environments. Although state-

Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy

Model ReleasesDGX agent

arXiv:2606.03142v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) show strong visualization interpretation, yet it is unclear whether their responses reflect genuine reasoning over

DMF: A Deterministic Memory Framework for Conversational AI Agents

Model ReleasesDGX agent

arXiv:2606.03463v1 Announce Type: new Abstract: Conversational AI agents require memory systems that are both scalable and semantically coherent across long interaction horizons. Existing approaches r

Do Value Vectors in Deep Layers Need Context from the Residual Stream?

Model ReleasesDGX agent

arXiv:2606.02780v1 Announce Type: new Abstract: The success of the transformer architecture as the backbone of modern LLMs is in large part due to its use of attention layers. An attention layer follo

Don't Forget Your Embeddings: Robust Knowledge Erasure via Precise Editing of Embeddings

Model ReleasesDGX agent

arXiv:2606.03695v1 Announce Type: new Abstract: As language models are increasingly deployed in real-world applications, the ability to erase specific knowledge from them becomes critical for safety a

Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation

Model ReleasesDGX agent

arXiv:2604.17220v2 Announce Type: replace-cross Abstract: Modeling coordination among generative agents in complex multi-round decision-making presents a core challenge for AI and operations managemen

Echelon: Auditable Aggregate-Only Language-Model Adaptation Across Privacy Boundaries

Model ReleasesDGX agent

arXiv:2606.02958v1 Announce Type: cross Abstract: Cross-organization language-model adaptation increasingly faces hard governance constraints: in many deployments, device-level model state-parameters,

Efficient ASR Training with Conversations that Never Happened

Model ReleasesDGX agent

arXiv:2606.03957v1 Announce Type: cross Abstract: Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose

🇸🇻 El Salvador now has its own open persona dataset Today, working with NVIDIA and WideLabs, a Latin American leader in sovereign AI, we h…

Model ReleasesDGX agent

🇸🇻 El Salvador now has its own open persona dataset Today, working with NVIDIA and WideLabs, a Latin American leader in sovereign AI, we have Nemotron-Personas-El-Salvador. It’s the first open dataset

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching

Model ReleasesDGX agent

arXiv:2606.03577v1 Announce Type: new Abstract: Wide-baseline matching (WBM) requires integrating geometric understanding, viewpoint changes, fine-grained perception, and occlusion reasoning, making i

eMEM: A Hybrid Spatio-Temporal Memory System For Embodied Agents

Model ReleasesDGX agent

arXiv:2606.03374v1 Announce Type: new Abstract: We present eMEM (Embodied Memory), a hybrid graph-based memory system for embodied agents operating in physical environments. Current agent memory archi

Enginuity: A Dataset and Benchmark for Vision-Language Understanding of Engineering Diagrams

Model ReleasesDGX agent

arXiv:2606.03410v1 Announce Type: new Abstract: Engineering diagrams pose a distinct challenge for vision-language models: unlike natural images or general documents, they encode information through d

EntSQL: A Benchmark for Grounding Text-to-SQL in Long-Context Enterprise Knowledge

Model ReleasesDGX agent

arXiv:2606.03363v1 Announce Type: new Abstract: Text-to-SQL enables natural language access to databases, and recent LLMs have substantially advanced its capabilities. Existing benchmarks such as Spid

ERP-XTTN: Interpretable Prototype-Guided Cross-Attention for Cross-Subject ERP Classification

Model ReleasesDGX agent

arXiv:2606.02939v1 Announce Type: new Abstract: Interpretable brain-computer interface classifiers that generalize across subjects without calibration remain an open challenge. We test whether prototy

EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction

Model ReleasesDGX agent

arXiv:2606.02971v1 Announce Type: new Abstract: Extracting reporting obligations from EU legislation is critical for assessing and reducing regulatory reporting burden. However, distinguishing reporti

Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers

Model ReleasesDGX agent

arXiv:2602.07842v2 Announce Type: replace Abstract: Confidence calibration is essential for making large language models (LLMs) reliable, yet existing training-free methods have been primarily studied

Evaluating LLMs' Effectiveness on Real-World Consumer Device Repair Questions

Model ReleasesDGX agent

arXiv:2606.03331v1 Announce Type: cross Abstract: Consumer device repair is an important but underexplored testbed for large language models (LLMs). Repair tasks require reasoning over incomplete prob

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

Model ReleasesDGX agent

arXiv:2606.03678v1 Announce Type: new Abstract: Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversa

Expanding Comfy With Claude Code https://x.com/i/broadcasts/1nJOLLXbrPWxR

Model ReleasesDGX agent

This X broadcast from ComfyUI likely discusses updates or new features for integrating Claude AI code capabilities with ComfyUI, a node-based UI for AI image generation and processing workflows. The s

Experience-Driven Dynamic Exits for LLMs with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.03113v1 Announce Type: new Abstract: Large Language Models suffer from slow autoregressive inference. While self-speculative decoding accelerates this process, its efficiency is hampered by

FinStressTS: A Parametric Synthetic Benchmark for Time-Series Forecasting in Finance

Model ReleasesDGX agent

arXiv:2606.03184v1 Announce Type: cross Abstract: Financial forecasting is difficult due to low signal-to-noise ratios, latent factors, heavy tails, regime shifts, and jumps. Real-world benchmarks off

Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling

Model ReleasesDGX agent

arXiv:2606.02837v1 Announce Type: cross Abstract: Accurate translation from Natural Language to First-Order Logic (NL-to-FOL) underpins neurosymbolic AI systems and Natural Language Inference (NLI), m

FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs

Model ReleasesDGX agent

arXiv:2506.01969v3 Announce Type: replace-cross Abstract: Efficient inference of Multi-Head Latent Attention (MLA) is challenged by deploying the DeepSeek-R1 671B model on a single Multi-GPU server. T

Forecasting Conceptual Diffusion in Science: The Case of Quantum Computing

Model ReleasesDGX agent

arXiv:2606.03919v1 Announce Type: cross Abstract: Understanding and anticipating scientific change requires models that distinguish between endogenous consolidation and exogenous diffusion of scientif

From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2606.03660v1 Announce Type: new Abstract: Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical

From Local Training to Large-Scale Mapping: A Comparative Assessment of Machine Learning and Deep Learning for Transferable Satellite-Derived Bathymetry

Model ReleasesDGX agent

arXiv:2606.02764v1 Announce Type: new Abstract: Satellite-derived bathymetry (SDB) from multispectral imagery is cost-effective but scales poorly across regions, especially in optically complex coasta

From Long News to Accurate Forecast: Importance-Aware Fusion and PRM-Guided Reflection for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2606.03097v1 Announce Type: new Abstract: Incorporating news into time series forecasting is appealing because news can reveal abrupt exogenous events that historical values alone cannot recover

From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds

Model ReleasesDGX agent

arXiv:2606.03557v1 Announce Type: new Abstract: As generative AI capabilities expand, AI-driven virtual worlds face a growing architectural challenge. Users interact through in-world interfaces in mul

From Script to Semantics: Prompting Strategies for African NLI

Model ReleasesDGX agent

arXiv:2606.03304v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains

Frontier models are powerful advisors. On @harvey's Legal Agent Benchmark, a GLM 5.1 worker using Claude Opus 4.7 as a sparse advisor reache…

Model ReleasesDGX agent

Frontier models are powerful advisors. On @harvey's Legal Agent Benchmark, a GLM 5.1 worker using Claude Opus 4.7 as a sparse advisor reached 18/100 all-pass versus 14/100 for Opus alone, at 39% of th

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models

Model ReleasesDGX agent

arXiv:2606.03165v1 Announce Type: cross Abstract: The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific Englis

G^2C-MT: Graph-Guided Context Selection for Document-Level Machine Translation

Model ReleasesDGX agent

arXiv:2606.03078v1 Announce Type: new Abstract: Effective document-level machine translation (DocMT) requires capturing long-range discourse dependencies. Recent work has explored retrieval-based and

Gate AI: LLM Security Benchmark Evaluation Methodology and Results

Model ReleasesDGX agent

arXiv:2606.02959v1 Announce Type: new Abstract: Published evaluations of prompt-injection and jailbreak detectors for Large Language Models often suffer from two systematic weaknesses: per-dataset thr

Gemma 4 12B is here! Dense, mid-sized Gemma that fits right on your laptop - released by @google under Apache 2.0 Available now in LM Studio…

Model ReleasesDGX agent

Gemma 4 12B is a newly released dense language model from Google available under the Apache 2.0 open license, designed for mid-sized computing resources that can run locally on personal computers. The

Gemma 4 model load issues fixed in engine version 2.20.1. lms runtime update --all

Model ReleasesDGX agent

Gemma 4 model load issues fixed in engine version 2.20.1. lms runtime update --all Gemma 4 12B is here! Dense, mid-sized Gemma that fits right on your laptop - released by @google under Apache 2.0 Ava

Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency

Model ReleasesDGX agent

arXiv:2606.03641v1 Announce Type: new Abstract: We investigate whether large language models produce different medical triage recommendations for identical neurological symptoms when only the patient'

Generating Rectifiable Measures through Neural Networks

Model ReleasesDGX agent

arXiv:2412.05109v2 Announce Type: replace Abstract: We derive universal approximation results for the class of (countably) m-rectifiable measures. Specifically, we prove that m-rectifiable measures ca

Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

Model ReleasesDGX agent

arXiv:2510.21011v3 Announce Type: replace-cross Abstract: As generative AI tools are increasingly used to portray people in professional roles, understanding their racial and gender representational b

GeoDrive-Bench: Benchmarking Region-Specific Multimodal Reasoning in Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.02774v1 Announce Type: new Abstract: Vision-language models (VLMs) for autonomous driving have shown promising performance, but their ability to handle region-specific traffic rules remains

Geometry-Aware Tabular Diffusion

Model ReleasesDGX agent

arXiv:2606.02607v1 Announce Type: cross Abstract: Tabular synthesis is critical for privacy-preserving sharing and augmentation, yet diffusion models rely on implicit mechanisms to capture inter-colum

Giving a talk this Sunday June 7th at @agihouse_org talking about DeepSeek V4 on @togethercompute! Come hang and let's talk inference!

Model ReleasesDGX agent

A Together AI representative announced a talk scheduled for Sunday, June 7th at AGI House focusing on DeepSeek V4 and inference optimization on Together Compute's platform. The event was promoted as a

GN0: Toward a Unified Paradigm for Generation, Evaluation, and Policy Learning in Visual-Language Navigation

Model ReleasesDGX agent

arXiv:2606.03682v1 Announce Type: new Abstract: Embodied navigation connects intelligent agents with the physical world and is fundamental for general robotic intelligence. Limited availability and qu

Google introduces Gemma 4 12B, a unified, encoder-free open multimodal model that can run locally on devices with 16GB of VRAM or unified memory (Carl Franzen/VentureBeat)

Model ReleasesDGX agent

Carl Franzen / VentureBeat: Google introduces Gemma 4 12B, a unified, encoder-free open multimodal model that can run locally on devices with 16GB of VRAM or unified memory — While many AI open source

Google launches Dreambeans, an AI app that curates daily stories from Google data

Model ReleasesDGX agent

Google LLC today launched Dreambeans, an experimental app from its Google Labs division that uses artificial intelligence to assemble a finite set of personalized daily stories drawn from a user’s own

.@GoogleDeepMind's Gemma 4 - 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes --model gemma4:1…

Model ReleasesDGX agent

.@GoogleDeepMind's Gemma 4 - 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes --model gemma4:12b-mlx Claude Code: ollama launch claude --model gemma4:12b-

Google's new Gemma 4 12B model is designed to run on any laptop with 16GB of RAM

Model ReleasesDGX agent

Google DeepMind released Gemma 4 12B, an open AI model that brings multimodal capabilities to everyday laptops by processing text, images, and audio natively without separate encoders. Small enough to

GPU-Parallel Multi-Task Reinforcement Learning with Demonstration Guided Policy Optimization

Model ReleasesDGX agent

arXiv:2606.03335v1 Announce Type: new Abstract: Large scale GPU-parallel reinforcement learning has changed what can be trained in robot simulation, yet most systems still optimize one specialist poli

Greener Than Humans? Environmental Attitudes in Large Language Models

Model ReleasesDGX agent

arXiv:2606.02741v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in sustainability-related decision support, reporting, and public communication, yet little systemati

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory

Model ReleasesDGX agent

arXiv:2606.03144v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning as

Had Claude Code build a snake game where the snake becomes aware it is in the game and then... stuff happens. Some impressive creative decis…

Model ReleasesDGX agent

Had Claude Code build a snake game where the snake becomes aware it is in the game and then... stuff happens. Some impressive creative decisions by the AI (& also some very AI ones), I just gave a fir

Hallucination Is Linearly Decodable from Mid-Layer Hidden States in Quantized LLMs

Model ReleasesDGX agent

arXiv:2606.02628v1 Announce Type: cross Abstract: We investigate whether open-source LLMs encode a linearly separable truthfulness signal in their hidden states, and at which network depth this signal

← Previous
1…174175176177178…377
Next →