AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,577 results
18 May 2026

Measuring Maximum Activations in Open Large Language Models

Model ReleasesDGX agent

arXiv:2605.15572v1 Announce Type: new Abstract: The dynamic range of activations is a first-order constraint for low-bit quantization, activation scaling, and stable LLM inference. Prior work characte

MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models

Model ReleasesDGX agent

arXiv:2605.15589v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in the mental health domain, yet it remains unclear how well they capture related biomedical knowledg

MI-CXR: A Benchmark for Longitudinal Reasoning over Multi-Interval Chest X-rays

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.15574v1 Announce Type: new Abstract: Longitudinal chest X-ray (CXR) interpretation requires reasoning over disease evolution across multiple patient visits, yet most existing medical VQA be

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

Model ReleasesDGX agent

arXiv:2605.15342v1 Announce Type: new Abstract: Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation

MorphoHELM: A Comprehensive Benchmark for Evaluating Representations for Microscopy-Based Morphology Assays

Model ReleasesDGX agent

arXiv:2605.15383v1 Announce Type: new Abstract: Microscopy images contain rich information about how cells respond to perturbations, making them essential to applications like drug screening. To quant

Multi-Fidelity Flow Matching: Cascaded Refinement of PDE Solutions

Model ReleasesDGX agent

arXiv:2605.16118v1 Announce Type: new Abstract: The source distribution in conditional flow matching is a design parameter that can be calibrated to data, not a default isotropic prior. We exploit thi

Multi-Probe Zero Collision Hash (MPZCH): Mitigating Embedding Collisions and Enhancing Model Freshness in Large-Scale Recommenders

Model ReleasesDGX agent

arXiv:2602.17050v3 Announce Type: replace Abstract: Embedding tables are critical components of large-scale recommendation systems, facilitating the efficient mapping of high-cardinality categorical f

MuteBench: Modality Unavailability Tolerance Evaluation for Incomplete Multimodal Fusion

Model ReleasesDGX agent

arXiv:2605.15235v1 Announce Type: new Abstract: Multimodal physiological data powers clinical AI systems from intensive care units to wearable devices, but sensors routinely fail in practice. Two fail

MyoChallenge 2025: A New Benchmark for Human Athletic Intelligence

Model ReleasesDGX agent

arXiv:2605.15650v1 Announce Type: new Abstract: Athletic performance represents the pinnacle of human motor intelligence, demanding rapid choices, precise control, agility, and coordinated physical ex

Navigating Potholes with Geometry-Aware Sharpness Minimization

Model ReleasesDGX agent

arXiv:2605.16134v1 Announce Type: cross Abstract: Sharpness-aware minimization (SAM) encourages flat minima by perturbing parameters along directions of high loss curvature, but treats all parameter d

Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters?

Model ReleasesDGX agent

arXiv:2507.10236v2 Announce Type: replace Abstract: As generative Artificial Intelligence (AI) advances, the realism of AI generated imagery has reached a threshold capable of deceiving even vigilant

NEW paper from Meta: Agentic Discovery of Neural Architectures. This is a hot new area of research! Keep an eye on it.

Model ReleasesDGX agent

NEW paper from Meta: Agentic Discovery of Neural Architectures. This is a hot new area of research! Keep an eye on it. NEW paper from Meta. (bookmark it) It's an agent system that autonomously discove

NEW paper worth reading. GPT-5.4 nano plus a critic-comparator orchestration loop hits 76.4% on SWE-bench Verified, matching standalone Gemi…

Model ReleasesDGX agent

NEW paper worth reading. GPT-5.4 nano plus a critic-comparator orchestration loop hits 76.4% on SWE-bench Verified, matching standalone Gemini 3 Pro and Claude Opus 4.5 Thinking. The trick is to selec

Njord: A Probabilistic Graph Neural Network for Ensemble Ocean Forecasting

Model ReleasesDGX agent

arXiv:2605.15470v1 Announce Type: new Abstract: Ocean dynamics are inherently chaotic, yet existing machine learning ocean models produce only deterministic forecasts. We introduce Njord, a probabilis

Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation

Model ReleasesDGX agent

arXiv:2605.16241v1 Announce Type: cross Abstract: Billion-parameter Vision-Language-Action (VLA) policies have recently shown impressive performance in robotic manipulation, yet their size and inferen

OgBench: A Framework for Evaluating Graph Neural Networks on Omics Data

Model ReleasesDGX agent

arXiv:2605.15511v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) have become the dominant framework for inductive graph-level learning. Yet most benchmarks focus on the regime n gg p, wher

okay this is going kinda viral and tbh my original text was kind of messy, so here's a second pass with the help of Claude: -- Implement <SP…

Model ReleasesDGX agent

okay this is going kinda viral and tbh my original text was kind of messy, so here's a second pass with the help of Claude: -- Implement <SPEC>. As you work maintain a running implementation-notes.htm

On RGB-TIR Stereo Calibration under Extreme Resolution Asymmetry

Model ReleasesDGX agent

arXiv:2605.15860v1 Announce Type: new Abstract: Accurate geometric calibration of RGB-thermal infrared (TIR) stereo camera systems is essential for multimodal building envelope analysis, yet remains c

One thing to watch for with Claude & GPT is that the models expose too much irrelevant history in their outputs. Slides are given footers sa…

Model ReleasesDGX agent

One thing to watch for with Claude & GPT is that the models expose too much irrelevant history in their outputs. Slides are given footers saying things like 'Better, more targeted version' if you aske

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints

Model ReleasesDGX agent

arXiv:2504.11320v3 Announce Type: replace-cross Abstract: Large language models now serve millions of users daily, with providers incurring costs exceeding $700,000 per day. Each request requires toke

PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control

Model ReleasesDGX agent

arXiv:2605.15963v1 Announce Type: new Abstract: Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet the

Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models

Model ReleasesDGX agent

arXiv:2509.22739v3 Announce Type: replace-cross Abstract: Language models (LMs) are typically post-trained for desired capabilities and behaviors via weight-based or prompt-based steering, but the for

PBT-Bench: Benchmarking AI Agents on Property-Based Testing

Model ReleasesDGX agent

arXiv:2605.15229v1 Announce Type: cross Abstract: Existing code benchmarks measure whether an agent can produce any test that reproduces a known bug, or whether it can produce a patch that fixes a des

PDRNN: Modular Data-driven Pedestrian Dead Reckoning on Loosely Coupled Radio- and Inertial-Signalstreams

Model ReleasesDGX agent

arXiv:2605.15252v1 Announce Type: cross Abstract: Modern pedestrian dead reckoning (PDR) systems rely on fusing noisy and biased estimates of position, velocity, and calibrated orientation derived fro

PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization

Model ReleasesDGX agent

arXiv:2605.15222v1 Announce Type: cross Abstract: Large language models (LLMs) can often generate functionally correct code, but their ability to produce efficient implementations for performance-crit

Perforated Neural Networks for Keyword Spotting

Model ReleasesDGX agent

arXiv:2605.15647v1 Announce Type: new Abstract: Edge machine learning presents a unique set of constraints not encountered in cloud-scale model deployment: strict memory budgets, limited compute, and

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models

Model ReleasesDGX agent

arXiv:2512.01843v2 Announce Type: replace Abstract: Driven by the growing capacity and training scale, Text-to-Video (T2V) generation models have recently achieved substantial progress in video qualit

Position: Early-Stage Quality Assurance in Annotation Pipelines Is More Cost-Effective Than Late-Stage Validation

Model ReleasesDGX agent

arXiv:2605.15714v1 Announce Type: cross Abstract: This position paper argues that the machine learning community should prioritize early-stage quality assurance in annotation pipelines over the prevai

Position: Ideas Should be the Center of Machine Learning Research

Model ReleasesDGX agent

arXiv:2605.15253v1 Announce Type: new Abstract: Machine learning research increasingly bifurcates into two disconnected modes: benchmark-driven engineering that prioritizes metrics over understanding,

Probabilistic Dating of Historical Manuscripts via Evidential Deep Regression on Visual Script Features

Model ReleasesDGX agent

arXiv:2605.06475v1 Announce Type: cross Abstract: We introduce a probabilistic approach for dating historical manuscript pages from visual features alone. Instead of aggregating centuries into classes

Prompting Amazon Nova 2 for content moderation

Model ReleasesDGX agent

In this post, you learn how to prompt Amazon Nova 2 Lite for content moderation using structured and free-form approaches, grounded in the MLCommons AILuminate Assessment Standard. The prompting techn

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels

Model ReleasesDGX agent

arXiv:2605.15208v1 Announce Type: cross Abstract: Large Language Models are routinely compressed via post-training quantization to reduce inference costs and memory footprint for cloud and edge deploy

Quantum Feature Pyramid Gating for Seismic Image Segmentation

Model ReleasesDGX agent

arXiv:2605.15370v1 Announce Type: cross Abstract: Accurate salt-body delineation is essential for seismic interpretation because salt structures distort wave propagation, complicate velocity-model bui

🚀🚀Qwen3.7 Preview lands on Arena ! Here come Qwen3.7-Max-Preview & Qwen3.7-Plus-Preview. Alibaba now #6 lab in Text, #5 in Vision.⚡️⚡️ Can…

Model ReleasesDGX agent

🚀🚀Qwen3.7 Preview lands on Arena ! Here come Qwen3.7-Max-Preview & Qwen3.7-Plus-Preview. Alibaba now #6 lab in Text, #5 in Vision.⚡️⚡️ Can't wait to release Qwen3.7 series models!Stay tuned! @arena Qw

RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning

Model ReleasesDGX agent

arXiv:2512.04457v2 Announce Type: replace Abstract: Removing specific data influence from large language models (LLMs) remains challenging, as retraining is costly and existing approximate unlearning

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition

Model ReleasesDGX agent

arXiv:2403.13805v2 Announce Type: replace-cross Abstract: CLIP (Contrastive Language-Image Pre-training) uses contrastive learning from noise image-text pairs to excel at recognizing a wide array of c

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation

Model ReleasesDGX agent

arXiv:2605.15239v1 Announce Type: new Abstract: Safety alignment often improves robustness to harmful queries at the cost of reasoning ability, a tradeoff known as the safety tax. A common cause is di

Registers Matter for Pixel-Space Diffusion Transformers

Model ReleasesDGX agent

arXiv:2605.16147v1 Announce Type: new Abstract: Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by exti

Reinforcement learning for adaptive interior point methods in convex quadratic programming

Model ReleasesDGX agent

arXiv:2509.07404v2 Announce Type: replace-cross Abstract: Quadratic programming is a workhorse of modern nonlinear optimization, control, and data science. Although regularized methods offer convergen

Representation Without Reward: A JEPA Audit for LLM Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.15394v1 Announce Type: cross Abstract: Joint-embedding predictive architectures (JEPAs) propose that a model should learn more useful abstractions when trained to predict latent representat

Retrieval-Augmented Large Language Models for Schema-Constrained Clinical Information Extraction

Model ReleasesDGX agent

arXiv:2605.15467v1 Announce Type: cross Abstract: Conversational nurse-patient transcripts contain actionable observations, but converting these transcripts into structured representations at scale re

RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades

Model ReleasesDGX agent

arXiv:2605.15846v1 Announce Type: cross Abstract: Coding agents are increasingly deployed in real software development, where a single version iteration requires months of coordinated work across many

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision

Model ReleasesDGX agent

arXiv:2605.15537v1 Announce Type: new Abstract: This paper introduces RTL-BenchMT, an agentic framework for dynamically maintaining RTL generation benchmarks. Large Language Models (LLMs) assisted aut

Rule2DRC: Benchmarking LLM Agents for DRC Script Synthesis with Execution-Guided Test Generation

Model ReleasesDGX agent

arXiv:2605.15669v1 Announce Type: new Abstract: Manufacturable chip layouts must satisfy thousands of geometry-based design rules, and design rule checking (DRC) enforces them by running executable DR

Run Claude Managed Agents with Vercel Sandbox

Model ReleasesDGX agent

This article describes how to run Claude's managed agents within Vercel's Sandbox environment, enabling developers to execute AI agent workloads on Vercel's infrastructure. The integration allows user

Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training

Model ReleasesDGX agent

arXiv:2605.16184v1 Announce Type: cross Abstract: Second-order methods offer an attractive path toward more sample-efficient LLM training, but their practical use is often blocked by the systems cost

SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?

Model ReleasesDGX agent

arXiv:2605.15777v1 Announce Type: new Abstract: Computer-Using Agents (CUAs) are rapidly extending large language models (LLMs) beyond text-based reasoning toward action execution in more complex envi

SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery

Model ReleasesDGX agent

arXiv:2510.22665v3 Announce Type: replace-cross Abstract: Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability. Although recent advances in self-

SDOF: Taming the Alignment Tax in Multi-Agent Orchestration with State-Constrained Dispatch

Model ReleasesDGX agent

arXiv:2605.15204v1 Announce Type: new Abstract: Multi-agent orchestration frameworks such as LangChain, LangGraph, and CrewAI route tasks through graph-based pipelines but do not enforce the stage con

Searching on a Budget: HW-NAS with 10 Latency Probes

Model ReleasesDGX agent

arXiv:2504.00663v2 Announce Type: replace Abstract: Existing hardware-aware NAS (HW-NAS) methods typically assume access to precise information circa the target device, either via analytical approxima

SemanticOpt: Towards LLM-Based Semantic Black-Box Optimization

Model ReleasesDGX agent

arXiv:2510.25404v3 Announce Type: replace-cross Abstract: Optimizing an experimental system can be extremely challenging when each experiment is expensive, time-consuming, or difficult to perform. Exi

SGR: A Stepwise Reasoning Framework for LLMs with External Subgraph Generation

Model ReleasesDGX agent

arXiv:2605.16117v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities across diverse NLP applications, such as translation, text generation, and question a

ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents

Model ReleasesDGX agent

arXiv:2605.16116v1 Announce Type: new Abstract: Developing and evaluating e-commerce web agents requires environments that preserve meaningful task structure while enabling controllable, reproducible,

SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces

Model ReleasesDGX agent

arXiv:2605.15215v1 Announce Type: new Abstract: Recently, skills have been widely adopted in large language model (LLM)-based agent systems across various domains. In existing frameworks, skills are t

SkyLink: A Large Vision-Language Model Driven Re-ranking Framework for Cross-View UAV geolocalization

Model ReleasesDGX agent

arXiv:2603.08063v3 Announce Type: replace Abstract: Cross-view UAV geolocalization is fundamentally a challenging large-scale image retrieval task, aiming to determine the geographic coordinates of Un

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning

Model ReleasesDGX agent

arXiv:2512.15693v2 Announce Type: replace Abstract: The misuse of AI-driven video generation technologies has raised serious social concerns, highlighting the urgent need for reliable AI-generated vid

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

Model ReleasesDGX agent

arXiv:2605.15710v1 Announce Type: new Abstract: Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use ev

Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models

Model ReleasesDGX agent

arXiv:2605.15424v1 Announce Type: new Abstract: Human trajectory forecasting is crucial for safe navigation in crowded environments, requiring models that balance accuracy with computational efficienc

SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval

Model ReleasesDGX agent

arXiv:2605.15868v1 Announce Type: new Abstract: In this work, we address the critical yet underexplored challenge of symmetric multimodal-to-multimodal (MM2MM) retrieval, where queries and contexts ar

Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning

Model ReleasesDGX agent

arXiv:2601.12894v2 Announce Type: replace-cross Abstract: Diffusion Policy has dominated action generation due to its strong capabilities for modeling multi-modal action distributions, but its multi-s

← Previous
1…243244245246247…377
Next →