AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,611 results
28 May 2026

On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note

Model ReleasesDGX agent

arXiv:2605.27563v1 Announce Type: cross Abstract: This short note presents a dimension-independent subgaussian concentration bound for Gaussian vectors under coordinate-wise nonlinear mappings. Discov

Optimal LTLf Synthesis

Model ReleasesDGX agent

arXiv:2605.11544v2 Announce Type: replace Abstract: Strategy synthesis typically follows an all-or-nothing paradigm, returning unrealisable whenever a specification cannot be guaranteed in an uncertai

Optimal ridge regularization revisited

Model ReleasesDGX agent

arXiv:2605.28679v1 Announce Type: new Abstract: We consider L^2-regularized linear (ridge) regression over a finite data sample X with bounded covariance and linear prediction targets y with additive


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Opus 4.8 formulated the hypotheses in advance, conducting data cleaning, did research on references, conducted analyses, did robustness chec…

Model ReleasesDGX agent

Opus 4.8 formulated the hypotheses in advance, conducting data cleaning, did research on references, conducted analyses, did robustness checks, and put out the whole paper in LaTEX style. GPT-5.5 foun

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents

Model ReleasesDGX agent

arXiv:2605.28158v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used to assist with operations research (OR) modeling, yet existing OR-oriented benchmarks often redu

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

Model ReleasesDGX agent

arXiv:2605.27378v1 Announce Type: new Abstract: Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have pro

Parameter-Efficient Generative Modeling with Controlled Vector Fields

Model ReleasesDGX agent

arXiv:2605.28267v1 Announce Type: new Abstract: We introduce a continuous-time generative modeling framework, motivated by the Chow-Rashevskii theorem, that builds expressive flows from a small set of

Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline

Model ReleasesDGX agent

arXiv:2605.27440v1 Announce Type: cross Abstract: Small changes to how a buyer phrases a question -- 'best CRM' vs 'top CRM' vs 'best CRM for a SaaS startup' -- produce substantially different brand r

Particle-Guided Diffusion Models for Partial Differential Equations

Model ReleasesDGX agent

arXiv:2601.23262v2 Announce Type: replace Abstract: We introduce a guided stochastic sampling method that augments sampling from diffusion models with physics-based guidance derived from partial diffe

PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI

Model ReleasesDGX agent

arXiv:2605.27545v1 Announce Type: new Abstract: Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe text

Patched-DeltaNet: Token-Level Event-Driven Memory for Linear-Time Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.27992v1 Announce Type: new Abstract: Time series anomaly detection is critical for maintaining the reliability of mission-critical systems. While Transformer-based models like PatchTST have

PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft

Model ReleasesDGX agent

arXiv:2605.27762v1 Announce Type: new Abstract: We present PEAM, a Parametric Embodied Agent Memory framework in Minecraft that transforms agent memory from inference-time retrieval into parameter-res

PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation

Model ReleasesDGX agent

arXiv:2601.18006v2 Announce Type: replace Abstract: We present PEAR (Pairwise Evaluation for Automatic Relative Scoring), a supervised quality estimation (QE) metric family that reframes reference-fre

PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective

Model ReleasesDGX agent

arXiv:2605.28819v1 Announce Type: cross Abstract: Parameter-efficient finetuning (PEFT) has become the standard approach for adapting large language models, yet evaluations largely emphasize downstrea

Periodic RoPE for Infinite Context LLMs

Model ReleasesDGX agent

arXiv:2605.27980v1 Announce Type: cross Abstract: The ability to process ultra-long contexts is crucial for large language models (LLMs) to perform long-horizon tasks. While recent efforts have extend

Personal Visual Memory from Explicit and Implicit Evidence

Model ReleasesDGX agent

arXiv:2605.28806v1 Announce Type: cross Abstract: Long-term memory is increasingly important for personalized AI agents, yet existing benchmarks and methods remain largely text-centric. Even when imag

Personalized Observation Normalization for Federated Reinforcement Learning in Simulation Environments with Heterogeneity

Model ReleasesDGX agent

arXiv:2605.27385v1 Announce Type: cross Abstract: Federated reinforcement learning (FedRL) enables multiple agents to collaboratively train a global policy without sharing raw data, making it ideal fo

Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models

Model ReleasesDGX agent

arXiv:2503.01829v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) demonstrate persuasive capabilities that rival human-level persuasion. While these capabilities can be used for s

PetroBench: A Benchmark for Large Language Models in Petroleum Engineering

Model ReleasesDGX agent

arXiv:2605.28032v1 Announce Type: new Abstract: Large Language Models are increasingly applied in the petroleum industry, highlighting the need for a domain-specific evaluation framework. This study d

PINE: Pruning Boosted Tree Ensembles with Conformal In-Distribution Prediction Equivalence

Model ReleasesDGX agent

arXiv:2605.28068v1 Announce Type: new Abstract: Tree ensembles are machine learning models with strong predictive performance and interpretability, and remain widely used for tabular data. Standard pr

Pittsburgh-based Gray Swan, which stress-tests AI models for top frontier AI labs, raised a 40M Series A at a 200M valuation co-led by Wing VC and Madrona (Rashi Shrivastava/Forbes)

Model ReleasesDGX agent

Rashi Shrivastava / Forbes: Pittsburgh-based Gray Swan, which stress-tests AI models for top frontier AI labs, raised a 40M Series A at a 200M valuation co-led by Wing VC and Madrona — Gray Swan works

Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

Model ReleasesDGX agent

arXiv:2605.28201v1 Announce Type: new Abstract: Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into ext

Plug-and-Play Benchmarking of Reinforcement Learning Algorithms for Large-Scale Flow Control

Model ReleasesDGX agent

arXiv:2601.15015v2 Announce Type: replace Abstract: Reinforcement learning (RL) has shown promising results in active flow control (AFC), yet progress in the field remains difficult to assess as exist

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2605.28237v1 Announce Type: cross Abstract: Real-world navigation is fundamentally driven by Points of Interest (POIs), yet reaching a precise POI remains a critical 'final-meters' challenge. Ex

PointQ-Bench: Benchmarking Diagnostic and Interpretable Point Cloud Quality Assessment

Model ReleasesDGX agent

arXiv:2605.28241v1 Announce Type: new Abstract: Point cloud quality plays a critical role in 3D acquisition, reconstruction, rendering, and perception, yet existing point cloud quality assessment (PCQ

PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

Model ReleasesDGX agent

arXiv:2605.27887v1 Announce Type: new Abstract: LLMs have shown strong performance across diverse financial tasks, yet portfolio management (PM), a critical financial decision-making task, remains poo

Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

Model ReleasesDGX agent

arXiv:2605.27958v1 Announce Type: cross Abstract: Linear probes trained on LLM activations are increasingly proposed as deception-detection metrics, yet report AUROC exceeding 0.96 on clean benchmarks

PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature

Model ReleasesDGX agent

arXiv:2605.28375v1 Announce Type: new Abstract: Prion diseases are rare, rapidly progressive, and fatal neurodegenerative disorders that remain difficult to diagnose, particularly in their early stage

Privately Estimating Monotone Statistics in Polynomial Time

Model ReleasesDGX agent

arXiv:2605.27912v1 Announce Type: cross Abstract: We study efficient differentially private algorithms for estimating monotone statistics, i.e., statistics that are monotone under the addition of new

Probabilistic Data-Driven Modelling of Astrophysical Transients: The Neural Process Family for Ultrafast and Class-Agnostic Light Curve Reconstruction with NightLANP

Model ReleasesDGX agent

arXiv:2605.27527v1 Announce Type: cross Abstract: Astrophysical observations taken from Earth are subject to weather, environmental, and scientific constraints that lead to sparse, irregular light cur

Probing for Knowledge Attribution in Large Language Models

Model ReleasesDGX agent

arXiv:2602.22787v2 Announce Type: replace-cross Abstract: Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, w

ProgVLA: Progress-Aware Robot Manipulation Skill Learning

Model ReleasesDGX agent

arXiv:2605.28231v1 Announce Type: cross Abstract: We present ProgVLA, a compact vision-language-action (VLA) model designed for reliable robot manipulation under tight compute and memory budgets. The

Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit

Model ReleasesDGX agent

arXiv:2605.27439v1 Announce Type: cross Abstract: AI assistants like ChatGPT and Claude are recommendation engines, not search engines: they answer commercial queries by directly nominating brands rat

Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement

Model ReleasesDGX agent

arXiv:2605.28360v1 Announce Type: new Abstract: Automatic prompt optimization (APO) has driven significant gains in LLM-based agentic workflows. However, existing methods treat each task's prompt as a

PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting

Model ReleasesDGX agent

arXiv:2605.28066v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated remarkable efficacy in text embedding, yet current adaptation methods like LoRA face significant bottle

Prompting Is All You Need: Multi-view Prompting Large Language Models for Aspect-Based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2605.28058v1 Announce Type: new Abstract: Recent work explored the capabilities of Large Language Models (LLMs) in Aspect-Based Sentiment Analysis (ABSA) through few-shot prompting, requiring su

ProvMind: Provenance-grounded reasoning for materials synthesis

Model ReleasesDGX agent

arXiv:2605.28487v1 Announce Type: new Abstract: Materials process optimization requires reasoning over routes, conditions, tools and causal dependencies, yet most computational formulations flatten sy

PrunePath: Towards Highly Structured Sparse Language Models

Model ReleasesDGX agent

arXiv:2605.28283v1 Announce Type: cross Abstract: Feed-forward networks (FFNs) dominate the parameter count and computation of modern language models, yet existing pruning methods often struggle to co

Pruning and Distilling Mixture-of-Experts into Dense Language Models

Model ReleasesDGX agent

arXiv:2605.28207v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory,

PubMedCausal: A Span-Level Annotated Corpus for Causal Relation Extraction in Biomedical Text

Model ReleasesDGX agent

arXiv:2605.28363v1 Announce Type: new Abstract: Causal relation extraction (CRE) is central to biomedical text mining, but current resources often conflate causal relations with broader associations,

Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

Model ReleasesDGX agent

arXiv:2605.28091v1 Announce Type: new Abstract: Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple

📢Qwen3.7-Max just hit #3 on ITbench-AA — a fresh benchmark testing how well models handle real-world enterprise IT tasks, agentic-style. 🔧…

Model ReleasesDGX agent

📢Qwen3.7-Max just hit #3 on ITbench-AA — a fresh benchmark testing how well models handle real-world enterprise IT tasks, agentic-style. 🔧Agentic era, go with Qwen.🏃🏃 Artificial Analysis and IBM Resea

RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration

Model ReleasesDGX agent

arXiv:2508.09449v2 Announce Type: replace Abstract: Reference-based Super Resolution (RefSR) improves upon Single Image Super Resolution (SISR) by leveraging high-quality reference images to enhance t

REED: Post-Training Representation Editing for Cross-Domain Linguistic Steganalysis

Model ReleasesDGX agent

arXiv:2605.28298v1 Announce Type: new Abstract: In real-world scenarios of linguistic steganalysis, tested texts usually come from unseen domains with different vocabularies, topics, writing styles, a

Reevaluating Policy Gradient Methods for Imperfect-Information Games

Model ReleasesDGX agent

arXiv:2502.08938v4 Announce Type: replace Abstract: In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information game

Reflective Dialogue between Teacher and Solver Agents for Video Question Answering

Model ReleasesDGX agent

arXiv:2605.27885v1 Announce Type: new Abstract: Various approaches have been proposed to adapt Vision-Language Models (VLMs) to specialized domains for Video Question Answering, including fine-tuning

ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing

Model ReleasesDGX agent

arXiv:2511.14584v3 Announce Type: replace-cross Abstract: We present ReflexGrad, a dual-process architecture for within-episode failure recovery in LLM agents without demonstrations. When agents commi

Regression Language Models for Code

Model ReleasesDGX agent

arXiv:2509.26476v2 Announce Type: replace-cross Abstract: We study code-to-metric regression: predicting numeric outcomes of code executions, a challenging task due to the open-ended nature of program

Relational Semantic Reasoning on 3D Scene Graphs for Open World Interactive Object Search

Model ReleasesDGX agent

arXiv:2603.05642v2 Announce Type: replace-cross Abstract: Open-world interactive object search in household environments requires understanding semantic relationships between objects and their surroun

Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

Model ReleasesDGX agent

arXiv:2605.28044v1 Announce Type: new Abstract: Cited RAG evaluation often treats visible sources as a grounding signal, but a real, topically relevant citation can still under-warrant the attached wo

ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions

Model ReleasesDGX agent

arXiv:2605.27819v1 Announce Type: cross Abstract: Sparse autoencoders are usually trained one layer at a time, even though transformer residual stream activations are strongly coupled across depth. Th

Resolution-free neural surrogates for geometric parameterization and mapping with spatially varying fields

Model ReleasesDGX agent

arXiv:2605.28551v1 Announce Type: new Abstract: Many imaging problems require computing spatial transformations induced by spatially varying intensity, feature, or density fields. Canonical examples i

Resource-Constrained Affect Modelling via Variance Regularisation Pruning

Model ReleasesDGX agent

arXiv:2605.27479v1 Announce Type: cross Abstract: Affective computing systems are increasingly embedded in pervasive and interactive environments, such as adaptive games, assistive technologies, and r

Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification

Model ReleasesDGX agent

arXiv:2512.12887v3 Announce Type: replace Abstract: 3D medical image classification is essential for modern clinical workflows. Medical foundation models (FMs) have emerged as a promising approach for

Revisiting Metafeatures to Explain Model Differences on Tabular Data

Model ReleasesDGX agent

arXiv:2605.28418v1 Announce Type: new Abstract: With the rise of tabular foundation models alongside traditional models still performing well on many tasks, choosing the right model for a tabular data

RGC: a radio AGN classifier based on deep learning. I. A semi-supervised multiclass model for VLA images

Model ReleasesDGX agent

arXiv:2510.22190v2 Announce Type: replace-cross Abstract: Bent radio active galactic nuclei (RAGNs) -- wide-angle tails (WATs) and narrow-angle tails (NATs) -- trace dense environments in galaxy group

RMPL: Relation-aware Multi-task Progressive Learning with Stage-wise Training for Multimedia Event Extraction

Model ReleasesDGX agent

arXiv:2602.13748v2 Announce Type: replace Abstract: Multimedia Event Extraction (MEE) aims to identify events and their arguments from documents that contain both text and images. It requires groundin

Robust Moment-Based Estimation via Spectral Gradient Reweighting

Model ReleasesDGX agent

arXiv:2605.27718v1 Announce Type: cross Abstract: Moment-based estimation is a theoretically attractive approach to parametric inference, especially when likelihood-based estimation is unavailable, mi

RW-TTT: Batched Serving for Request-Owned Test-Time Training State

Model ReleasesDGX agent

arXiv:2605.28053v1 Announce Type: new Abstract: Test-time training (TTT) adapts an LLM during generation by reading and updating request-owned state, such as fast weights, low-rank deltas, or streamin

Safe In-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2509.25582v3 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks w

← Previous
1…202203204205206…377
Next →