AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens

DGX agent

arXiv:2512.17375v2 Announce Type: replace-cross Abstract: LLM-as-a-Judge systems supply the reward signal in modern RLHF and RLVR pipelines, but their binary verdict reduces to a single linear readout

model-releasesarxiv-cs-cl
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning

DGX agent

arXiv:2605.28192v1 Announce Type: new Abstract: Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across b

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Agentic Separation Logic Specification Synthesis

DGX agent

arXiv:2605.27531v1 Announce Type: cross Abstract: Specification synthesis, the task of automatically inferring formal specifications from program implementations and natural language, is important for

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Aligning Language Model Benchmarks with Pairwise Preferences

DGX agent

arXiv:2602.02898v2 Announce Type: replace Abstract: Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find that bench

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models

DGX agent

arXiv:2602.18481v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has led to a surge of financial benchmarks, evolving from static knowledge evaluation to

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

AlphaTransit: Learning to Design City-scale Transit Routes

DGX agent

arXiv:2605.28730v1 Announce Type: new Abstract: Designing a transit network requires many sequential route extension decisions, but their quality is often visible only after the full network is assemb

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

An Enhanced Large Neighborhood Search Approach for the Capacitated Facility Location Problem with Incompatible Customers

DGX agent

arXiv:2605.28337v1 Announce Type: new Abstract: A new variant of the classic capacitated facility location problem, which considers incompatibilities between customers, has recently been introduced in

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Analyzing Quality-Latency-Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation

DGX agent

arXiv:2605.28222v1 Announce Type: new Abstract: We study quality-latency-resource trade-offs in a documentation-grounded retrieval-augmented generation (RAG) system that uses Low-Rank Adaptation (LoRA

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

DGX agent

arXiv:2605.27761v1 Announce Type: new Abstract: The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Apple Intelligence Foundation Language Models

DGX agent

arXiv:2407.21075v2 Announce Type: replace Abstract: We present foundation language models developed to power Apple Intelligence features, including a ~3 billion parameter model designed to run efficie

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?

DGX agent

arXiv:2508.11011v2 Announce Type: replace Abstract: Construction safety inspections typically involve a human inspector identifying safety concerns on-site. With the rise of powerful Vision Language M

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Argument Quality Assessment with Large Language Models: A Pairwise Bradley-Terry Approach

DGX agent

arXiv:2605.28313v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in tasks related to reasoning and judgment. However, assessing the quality of arg

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers

DGX agent

arXiv:2512.09800v2 Announce Type: replace Abstract: Low-power microcontroller (MCU) hardware is currently evolving from single-core architectures to predominantly multi-core architectures. In parallel

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents

DGX agent

arXiv:2605.28108v1 Announce Type: new Abstract: A long-lived LLM agent, such as OpenClaw, earns its value by acting on a user's preferences and constraints across sessions, not just the current reques

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications

DGX agent

arXiv:2605.27472v1 Announce Type: cross Abstract: Assertion-based verification (ABV) is a cornerstone of modern hardware design, yet manually translating design intent into formal SystemVerilog Assert

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Assessing Factual Music Comprehension in Large Audio Language Models

DGX agent

arXiv:2511.05550v2 Announce Type: replace-cross Abstract: Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference

DGX agent

arXiv:2505.19342v2 Announce Type: replace-cross Abstract: Multi-device inference can reduce Transformer latency by parallelizing computation. However, existing methods require high inter-device bandwi

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Asynchronous Remote Sensing Time-Series Fusion for Cloud Removal and Anytime Reconstruction

DGX agent

arXiv:2605.27726v1 Announce Type: new Abstract: Frequent cloud cover severely limits the usability of Sentinel-2 (S2) optical time series for Earth surface monitoring. Sentinel-1 (S1) SAR provides all

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios

DGX agent

arXiv:2605.27995v1 Announce Type: new Abstract: Large language model (LLM)-based agents have shown strong capabilities in using external tools to solve complex tasks. However, existing evaluations oft

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

ATLAS: All-round Testing of Long-context Abilities across Scales

DGX agent

arXiv:2605.28079v1 Announce Type: new Abstract: Long-context language models now advertise context windows up to millions of tokens, yet evaluations typically report a single length or a narrow task f

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity

DGX agent

arXiv:2605.28640v1 Announce Type: new Abstract: Efficient inference is critical for long-context language models, where attention computation and KV-cache access dominate the cost. Recent work RAT+, i

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Automating Formal Verification with Agent-Guided Tree Search

DGX agent

arXiv:2605.27485v1 Announce Type: cross Abstract: Formal verification offers a path to provably correct software, but writing verified code remains expensive enough that the technique is rarely used i

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

DGX agent

arXiv:2605.28642v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment par

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Bayesian Optimization Parameter Tuning Framework for a Lyapunov Based Path Following Controller

DGX agent

arXiv:2512.12649v2 Announce Type: replace Abstract: Parameter tuning in real-world experiments is constrained by the limited evaluation budget available on hardware. The path-following controller stud

model-releasesarxiv-cs-ro
28 May 2026
Model Releases

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

DGX agent

arXiv:2605.28508v1 Announce Type: new Abstract: Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape us

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment

DGX agent

arXiv:2604.00913v2 Announce Type: replace-cross Abstract: 2D assembly diagrams are often abstract and hard to follow, creating a need for intelligent assistants that can monitor progress, detect error

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Benchmarking Fairness in Spiking Neural Networks: Data Bias, Spurious Features, and Hardware Effects

DGX agent

arXiv:2605.27407v1 Announce Type: cross Abstract: Evaluating fairness in Spiking Neural Networks (SNNs) demands rigorous benchmarks that reflect real-world complexities, yet existing assessments remai

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Benchmarking Inductive Biases for Multivariate Time-Series Anomaly Detection with a Robust Multi-View Channel-Graph Detector

DGX agent

arXiv:2605.28103v1 Announce Type: new Abstract: We present a unified experiment, analysis, and benchmark study of multivariate time-series (MTS) anomaly detection. Ten family-representative detectors

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Benchmarking Ultrasound Foundation Models for Fetal Plane Classification

DGX agent

arXiv:2605.27796v1 Announce Type: cross Abstract: Ultrasound is widely used in obstetric care due to its safety, accessibility, and real-time imaging. However, interpretation remains operator-dependen

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems

DGX agent

arXiv:2605.27492v1 Announce Type: cross Abstract: LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

DGX agent

arXiv:2605.28183v1 Announce Type: cross Abstract: We introduce the BenGER (Benchmark for German Law) dataset for evaluating LLM systems on subsumption-based legal reasoning in German law. The BenGER d

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

DGX agent

arXiv:2605.28301v1 Announce Type: new Abstract: Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond Binary Moral Judgment: Modeling Ethical Pluralism in AI

DGX agent

arXiv:2605.28707v1 Announce Type: new Abstract: Critical decision-making in socially consequential spaces is increasingly involving AI systems at varying capacities. Yet, despite the ubiquity of auton

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

DGX agent

arXiv:2502.05242v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain uncl

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

DGX agent

arXiv:2509.23074v3 Announce Type: replace-cross Abstract: In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark lea

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond Motion Primitives: Behavioral Activity Recognition from Head-Mounted IMU

DGX agent

arXiv:2605.27464v1 Announce Type: cross Abstract: AR smart glasses need continuous behavioral context to offer proactive assistance, yet their most practical always-on sensor, the head-mounted Inertia

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

DGX agent

arXiv:2605.28465v1 Announce Type: new Abstract: Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generation

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Bias Leaves a Gradient Trail: Label-Free Bias Identification via Gradient Probes on Concept Decompositions

DGX agent

arXiv:2605.28780v1 Announce Type: new Abstract: Vision classifiers can exploit spurious correlations, achieving high in-distribution accuracy yet failing under distribution shift. Existing approaches

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Bilinear Coordinate Alignment for Training-Free Task-Vector Transfer

DGX agent

arXiv:2605.28444v1 Announce Type: new Abstract: Fine-tuning large-scale pre-trained models is a recent prevalent paradigm for adapting general representations to specialized tasks. However, when a new

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

BioELX: Cross-lingual Biomedical Entity Linking via Alias-based Retrieval and LLM Ranking

DGX agent

arXiv:2605.27380v1 Announce Type: cross Abstract: Cross-lingual biomedical entity linking (BEL) maps mentions in any language to unique identifiers in a biomedical knowledge base (KB), supporting clin

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

BlazeEdit: Generalist Image Editing on Mobile Devices with Image-to-Image Diffusion Models

DGX agent

arXiv:2605.28067v1 Announce Type: new Abstract: The remarkable generation quality of modern diffusion models often comes at the cost of massive parameter counts, which necessitate server-side inferenc

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Bounded-Compute Multimodal Regression for Product-Rating Prediction

DGX agent

arXiv:2605.27737v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generatio

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

DGX agent

arXiv:2605.27383v1 Announce Type: cross Abstract: Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However,

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

BuddyBench: A Privacy-Constrained Multi-Task Benchmark for Pediatric Social-Communication Personalization

DGX agent

arXiv:2605.28089v1 Announce Type: new Abstract: BuddyBench introduces a privacy-constrained multi-task benchmark for pediatric social-communication personalization. Unlike existing neurodevelopmental

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Building Community-Centred NLP Resources for Puno Quechua

DGX agent

arXiv:2605.28253v1 Announce Type: new Abstract: The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

DGX agent

arXiv:2510.05291v2 Announce Type: replace Abstract: As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensitivity to culturally diverse entities becomes increasingly im

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Can Decision Trees Teach Large Language Models? Distilling Verbalized Knowledge for Molecular Property Prediction

DGX agent

arXiv:2603.12344v2 Announce Type: replace Abstract: Molecular Property Prediction (MPP) is a fundamental problem in drug discovery that has recently attracted growing attention. Large Language Models

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Can Large Language Models Handle Discourse Particles? A Case Study of Colloquial Malay

DGX agent

arXiv:2605.28782v1 Announce Type: new Abstract: Discourse particles, such as extit{well} and extit{kind of}, are crucial components that enable LLMs to ``speak'' more like humans. They are used to con

model-releasesarxiv-cs-cl
28 May 2026
← Previous
1…185186187188189…361
Next →