AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,611 results
27 May 2026

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

Model ReleasesDGX agent

arXiv:2605.27235v1 Announce Type: new Abstract: Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, an

MTL-FNO: A Lightweight Multi-Task Fourier Neural Operator for Sparse Field Reconstruction

Model ReleasesDGX agent

arXiv:2605.26718v1 Announce Type: new Abstract: Efficient onboard multi-field sparse reconstruction is essential for the autonomous operation of aerospace vehicles. While existing deep learning models

Multi-Agent Causal Discovery Using Large Language Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2407.15073v4 Announce Type: replace Abstract: Causal discovery aims to identify causal relationships between variables and is a fundamental problem across the sciences. Traditional statistical c

MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding

Model ReleasesDGX agent

arXiv:2605.26320v1 Announce Type: cross Abstract: The application of generalist multimodal models (GMMs) to specialized scientific domains remains limited due to the scarcity of comprehensive domain-s

Near-Optimal Regret in Adversarial Kernel Bandits

Model ReleasesDGX agent

arXiv:2605.26585v1 Announce Type: new Abstract: We study the adversarial kernel bandit problem, in which the loss at each round is induced by an arbitrary bounded element of a reproducing kernel Hilbe

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

Model ReleasesDGX agent

arXiv:2605.26895v1 Announce Type: cross Abstract: Normalization layers in modern large language models (LLMs) consist of a deterministic normalization operation and a learnable scale vector. While the

NestedKV: Nested Memory Routing for Long-Context KV Cache Compression

Model ReleasesDGX agent

arXiv:2605.26678v1 Announce Type: new Abstract: Long-context language models are limited by the memory footprint of the key-value (KV) cache. Existing training-free KV compression methods usually rank

Neural Autoregressive Control Variates for the Quantum Monte Carlo Sign Problem

Model ReleasesDGX agent

arXiv:2605.26814v1 Announce Type: cross Abstract: We train a pair of autoregressive models to construct zero-mean control variates to mitigate the sign problem in quantum Monte Carlo simulations. The

Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation

Model ReleasesDGX agent

arXiv:2605.26844v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own rollouts with token-level teacher supervision. Recent selective OPD methods exploit the non-uni

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

Model ReleasesDGX agent

arXiv:2605.26584v1 Announce Type: new Abstract: Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

Model ReleasesDGX agent

arXiv:2505.17163v2 Announce Type: replace-cross Abstract: Recent advancements in multimodal slow-thinking systems have demonstrated remarkable performance across various visual reasoning tasks. Howeve

ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection

Model ReleasesDGX agent

arXiv:2508.01253v2 Announce Type: replace Abstract: Existing studies typically investigate domain shift and category shift as independent problems, however, in real-world scenarios, the two types of s

Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2603.16654v2 Announce Type: replace-cross Abstract: Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, espec

OMD-GraphRAG: Enhancing GraphRAG with Ontology-Guided Extraction, Multi-Dimensional Clustering and Dual-Channel Fusion

Model ReleasesDGX agent

arXiv:2603.25152v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems face significant challenges in complex reasoning, multi-hop queries, and domain-specific QA. While exis

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

Model ReleasesDGX agent

arXiv:2605.26485v1 Announce Type: cross Abstract: We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-vi

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation

Model ReleasesDGX agent

arXiv:2605.26641v1 Announce Type: new Abstract: Unified multimodal embedding spaces have become the standard interface for cross-modal retrieval and multimodal RAG, and recent audio-video-text (AVT) e

OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling

Model ReleasesDGX agent

arXiv:2605.26322v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to infer others' knowledge, intentions, and emotions, is commonly evaluated in large language models (LLMs) using end-

On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning

Model ReleasesDGX agent

arXiv:2605.27083v1 Announce Type: new Abstract: Counterfactual tuning (CFT) has emerged as a promising paradigm for Large Language Model (LLM) unlearning by training models to generate alternative fic

On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs

Model ReleasesDGX agent

arXiv:2510.05864v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly operate on long inputs, yet their behavior when harmful sentences are sparsely embedded within such inputs

ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research

Model ReleasesDGX agent

arXiv:2601.21008v3 Announce Type: replace-cross Abstract: Operations Research practitioners debug infeasible models through an iterative process: inspecting Irreducible Infeasible Subsystems ( IIS), i

OSMa-Bench++: Toward Open-Ended Benchmarking of Semantic Mapping for Manipulation with Prompt-Generated Synthetic Scenes

Model ReleasesDGX agent

arXiv:2605.26831v1 Announce Type: new Abstract: Semantic mapping methods are increasingly used as intermediate scene representations for downstream robotic reasoning and manipulation, yet their evalua

PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech

Model ReleasesDGX agent

arXiv:2605.26978v1 Announce Type: new Abstract: Text-to-speech (TTS) evaluation for low-resource non-Latin-script languages can fail when it relies on a single ASR round-trip word error rate (WER). A

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

Model ReleasesDGX agent

arXiv:2511.20586v4 Announce Type: replace Abstract: Trustworthiness has become a key requirement for the deployment of artificial intelligence systems in safety-critical applications. Conventional eva

Periodic Topological Deep Learning for Polymer Design and Discovery

Model ReleasesDGX agent

arXiv:2605.26833v1 Announce Type: cross Abstract: Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery challenging.

PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark

Model ReleasesDGX agent

arXiv:2506.00250v4 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing

PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions

Model ReleasesDGX agent

arXiv:2605.27015v1 Announce Type: new Abstract: Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We in

Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

Model ReleasesDGX agent

arXiv:2602.17003v2 Announce Type: replace-cross Abstract: Large language models have advanced web agents, yet current agents lack personalization capabilities. Since users rarely specify every detail

'PhyWorldBench': A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

Model ReleasesDGX agent

arXiv:2507.13428v3 Announce Type: replace-cross Abstract: Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurate

PIDM-DP: Physics-Informed Diffusion with Dormand-Prince Integration for Chaotic System Identification and State Reconstruction across Multiple Dynamical Regimes

Model ReleasesDGX agent

arXiv:2605.26619v1 Announce Type: new Abstract: Reconstructing continuous state trajectories of chaotic dynamical systems from sparse, noisy observations remains a fundamental open problem in nonlinea

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

Model ReleasesDGX agent

arXiv:2605.27258v1 Announce Type: cross Abstract: Building state-of-the-art text-to-speech (TTS) systems typically demands millions of hours of proprietary data and complex multi-stage architectures,

PitchBench: Measuring Pitch Hearing in Audio-Language Models

Model ReleasesDGX agent

arXiv:2605.26176v1 Announce Type: cross Abstract: Audio-language models (ALMs) are increasingly used in real-world applications that require understanding music, from music tutoring and transcription

Position: AI Safety Requires Effective Controllability

Model ReleasesDGX agent

arXiv:2605.27117v1 Announce Type: new Abstract: AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing ha

PRBench: A Standardized Probabilistic Robustness Benchmark

Model ReleasesDGX agent

arXiv:2511.01724v3 Announce Type: replace Abstract: Deep learning models are notoriously vulnerable to imperceptible perturbations. Most existing research centers on adversarial robustness (AR), which

Pretrained Approximators for Low-Thrust Trajectory Cost and Reachability

Model ReleasesDGX agent

arXiv:2605.26790v1 Announce Type: new Abstract: Low-thrust trajectory design relies heavily on repeated evaluations of fuel consumption and transfer feasibility, which require expensive optimal contro

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

Model ReleasesDGX agent

arXiv:2605.26730v1 Announce Type: new Abstract: The rapid growth in submissions to machine learning venues has strained the scientific peer-review system and intensified interest in LLM-based automate

PRISM: Position-encoded Regressive Inverse Spectral Model for Multilayer Thin-Film Design

Model ReleasesDGX agent

arXiv:2605.26502v1 Announce Type: new Abstract: The inverse problem of multilayer thin-film optical coatings design represents a complex combinatorial-continuous optimization challenge. We present PRI

Probing Cultural Awareness in LLMs: A Case Study of Cross-Culture Aesthetic Stylistics

Model ReleasesDGX agent

arXiv:2605.27296v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in diverse cultural contexts, yet their ability to master aesthetic stylistics, i.e., the strateg

Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals

Model ReleasesDGX agent

arXiv:2605.26999v1 Announce Type: new Abstract: Prompt injection poses a critical threat to the safe deployment of large language models, yet existing detection approaches are typically evaluated unde

Provably Communication-Efficient and Privacy-Preserving Federated Graph Neural Networks

Model ReleasesDGX agent

arXiv:2605.26243v1 Announce Type: new Abstract: Graph neural networks (GNNs) achieve strong performance on relational data, but real-world graphs are often distributed across organizations that cannot

Qiskit QuantumKatas: Adapting Microsoft's Quantum Computing exercises for LLM evaluation

Model ReleasesDGX agent

arXiv:2605.27210v1 Announce Type: cross Abstract: We adapt Microsoft's QuantumKatas -- a well-established quantum computing curriculum -- from Q# to Qiskit, the most widely-adopted quantum computing f

Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering

Model ReleasesDGX agent

arXiv:2605.27164v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems for question answering typically retrieve evidence by semantic similarity between the query and document ch

🚀🚀 Qwen3.7-Max just hit #4 on Code Arena, on par with Claude Opus 4.6 ,top-ranked Chinese lab on the board! @arena More to ship. Stay tune…

Model ReleasesDGX agent

🚀🚀 Qwen3.7-Max just hit #4 on Code Arena, on par with Claude Opus 4.6 ,top-ranked Chinese lab on the board! @arena More to ship. Stay tuned. 🕶️ Qwen3.7 Max (20250517) debuts at #4 in Code Arena: Front

RCSP: Risk-Sensitive Conjectural Scenario Planning for Safe Dynamic Robot Navigation

Model ReleasesDGX agent

arXiv:2605.26348v1 Announce Type: new Abstract: Mobile robots can fail before they collide: a velocity that is safe now may commit the robot to a passage that moving obstacles will soon close. We stud

Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions

Model ReleasesDGX agent

arXiv:2605.26414v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve impressive accuracy on mathematical reasoning benchmarks, yet their performance drops when problems are modified wi

Receipt Replay OOD: A Small Benchmark for Screen Replay Detection Under Domain Shift

Model ReleasesDGX agent

arXiv:2605.26855v1 Announce Type: new Abstract: Public datasets such as DLC-2021, SynID, and KID34K have significantly contributed to research on presentation attack detection for identity documents,

Reliable Extraction of Clinical Follow-Up Instructions: A Hybrid Neural-Symbolic Pipeline

Model ReleasesDGX agent

arXiv:2605.26560v1 Announce Type: cross Abstract: Objective. Outpatient notes carry follow-up instructions pairing actions with future times ('MRI brain in two weeks'). Extracting (action, date) pairs

ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference

Model ReleasesDGX agent

arXiv:2605.27081v1 Announce Type: cross Abstract: Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining hi

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure

Model ReleasesDGX agent

arXiv:2601.10566v5 Announce Type: replace Abstract: Entity-level unlearning is usually evaluated by what a model says: whether it stops naming the target, refuses a query, or shifts a Truth Ratio dist

Resolving Ambiguity in Composed Image Retrieval via Calibrated Interaction

Model ReleasesDGX agent

arXiv:2605.24634v2 Announce Type: replace Abstract: Composed image retrieval (CIR) searches a corpus with a reference image and a text describing how to modify it. Despite rapid progress from triplet-

Rethinking the Multilingual Reasoning Gap with Layer Swap

Model ReleasesDGX agent

arXiv:2605.26735v1 Announce Type: new Abstract: Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior wor

Risk Averse Alert Prioritization for IDS Using Subnormal Gaussian Fuzzy Models

Model ReleasesDGX agent

arXiv:2605.27299v1 Announce Type: cross Abstract: Modern intrusion detection systems generate thousands of alerts daily, but alert fatigue severely limits security operations effectiveness due to too

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data

Model ReleasesDGX agent

arXiv:2605.26971v1 Announce Type: new Abstract: The proliferation of Reinforcement Learning from Verifiable Rewards (RLVR) datasets has exacerbated provenance collapse due to unclear lineage among exi

RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction

Model ReleasesDGX agent

arXiv:2605.26862v1 Announce Type: new Abstract: Accurate road segmentation from aerial imagery is fundamental to many geospatial applications. However, existing datasets often suffer from limited scen

Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

Model ReleasesDGX agent

arXiv:2605.27134v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking,

Scaling GraphLLM with Bilevel-Optimized Sparse Querying

Model ReleasesDGX agent

arXiv:2602.09038v2 Announce Type: replace-cross Abstract: LLMs have recently shown strong potential in enhancing node-level tasks on text-attributed graphs (TAGs) by providing explanation features. Ho

Scheduled Style Injection: Expanding the Style-Content Pareto Frontier in Training-Free Diffusion-based Style Transfer

Model ReleasesDGX agent

arXiv:2605.26538v1 Announce Type: new Abstract: Style transfer with pre-trained diffusion models has advanced rapidly, but a core question remains underexplored: where in the model should style inject

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

Model ReleasesDGX agent

arXiv:2605.26340v1 Announce Type: new Abstract: Autonomous research agents produce competitive solutions and professional-looking manuscripts, yet their outputs contain verifiability failures undetect

SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge Graphs

Model ReleasesDGX agent

arXiv:2512.04868v2 Announce Type: replace-cross Abstract: Knowledge-based conversational question answering (KBCQA) confronts persistent challenges in resolving coreference, modeling contextual depend

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?

Model ReleasesDGX agent

arXiv:2605.26548v1 Announce Type: cross Abstract: Large language models (LLMs) now support automated software security tasks, including vulnerability discovery and proof-of-concept (PoC) generation. E

SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability

Model ReleasesDGX agent

arXiv:2605.26788v1 Announce Type: cross Abstract: Large language models (LLMs) achieve impressive performance when a task is fully specified in a single turn, yet the same models lose up to 39% of tha

← Previous
1…207208209210211…377
Next →