AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,280 results
15 Apr 2026

Orthogonal Subspace Projection for Continual Machine Unlearning via SVD-Based LoRA

Model ReleasesDGX agent

arXiv:2604.12526v1 Announce Type: cross Abstract: Continual machine unlearning aims to remove the influence of data that should no longer be retained, while preserving the usefulness of the model on e

Parametric Interpolation of Dynamic Mode Decomposition for Predicting Nonlinear Systems

Model ReleasesDGX agent

arXiv:2604.12103v1 Announce Type: cross Abstract: We present parameter-interpolated dynamic mode decomposition (piDMD), a parametric reduced-order modeling framework that embeds known parameter-affine

Parcae: Scaling Laws For Stable Looped Language Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.12946v1 Announce Type: new Abstract: Traditional fixed-depth architectures scale quality by increasing training FLOPs, typically through increased parameterization, at the expense of a high

ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving

Model ReleasesDGX agent

arXiv:2604.00136v2 Announce Type: replace-cross Abstract: Multi-model LLM serving operates in a non-stationary, noisy environment: providers revise pricing, model quality can shift or regress without

Parsing complex tables in PDFs is extremely challenging. Existing metrics for measuring table accuracy, like TEDS (tree edit distance simila…

Model ReleasesDGX agent

Parsing complex tables in PDFs is extremely challenging. Existing metrics for measuring table accuracy, like TEDS (tree edit distance similarity), overweight exact table structure and underweight sema

Physically Accurate Rigid-Body Dynamics in Particle-Based Simulation

Model ReleasesDGX agent

arXiv:2603.14634v3 Announce Type: replace Abstract: Robotics demands simulation that can reason about the diversity of real-world physical interactions, from rigid to deformable objects and fluids. Cu

Policy-Invisible Violations in LLM-Based Agents

Model ReleasesDGX agent

arXiv:2604.12177v1 Announce Type: new Abstract: LLM-based agents can execute actions that are syntactically valid, user-sanctioned, and semantically appropriate, yet still violate organizational polic

PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models

Model ReleasesDGX agent

arXiv:2604.12995v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly integrated into real-world decision-making, including in the domain of public policy. Yet, their ability t

Polynomial Expansion Rank Adaptation: Enhancing Low-Rank Fine-Tuning with High-Order Interactions

Model ReleasesDGX agent

arXiv:2604.11841v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) is a widely used strategy for efficient fine-tuning of large language models (LLMs), but its strictly linear structure fund

ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems

Model ReleasesDGX agent

arXiv:2604.11943v1 Announce Type: cross Abstract: An OS kernel that runs LLM inference internally can read logit distributions before any text is generated -- and act on them as a governance primitive

PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.12652v1 Announce Type: cross Abstract: Reinforcement learning (RL) can improve the prompt following capability of text-to-image (T2I) models, yet obtaining high-quality reward signals remai

Public Profile Matters: A Scalable Integrated Approach to Recommend Citations in the Wild

Model ReleasesDGX agent

arXiv:2603.17361v2 Announce Type: replace-cross Abstract: Proper citation of relevant literature is essential for contextualising and validating scientific contributions. While current citation recomm

Quantile Q-Learning: Revisiting Offline Extreme Q-Learning with Quantile Regression

Model ReleasesDGX agent

arXiv:2511.11973v2 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables policy learning from fixed datasets without further environment interaction, making it particularly valu

QuarkMedSearch: A Long-Horizon Deep Search Agent for Exploring Medical Intelligence

Model ReleasesDGX agent

arXiv:2604.12867v1 Announce Type: new Abstract: As agentic foundation models continue to evolve, how to further improve their performance in vertical domains has become an important challenge. To this

Ran ChatGPT Plus and Claude Pro side by side for 30 days, here's what I found as a daily ChatGPT user

Model ReleasesDGX agent

A Reddit post from r/ChatGPT in which a longtime ChatGPT Plus user shares findings after running both ChatGPT Plus and Claude Pro simultaneously for 30 days. The post likely reflects a real-world comp

RankOOD -- Class Ranking-based Out-of-Distribution Detection

Model ReleasesDGX agent

arXiv:2511.19996v2 Announce Type: replace Abstract: We propose RankOOD, a rank-based Out-of-Distribution (OOD) detection approach based on training a model with the Placket-Luce loss, which is now ext

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.12371v1 Announce Type: new Abstract: We study typographic prompt injection attacks on vision-language models (VLMs), where adversarial text is rendered as images to bypass safety mechanisms

ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance

Model ReleasesDGX agent

arXiv:2604.12378v1 Announce Type: new Abstract: Despite advances in multilingual capabilities, most large language models (LLMs) remain English-centric in their training and, crucially, in their produ

Red Teaming Large Reasoning Models

Model ReleasesDGX agent

arXiv:2512.00412v4 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) have emerged as a powerful advancement in multi-step reasoning tasks, offering enhanced transparency and logical

ReflectCAP: Detailed Image Captioning with Reflective Memory

Model ReleasesDGX agent

arXiv:2604.12357v1 Announce Type: new Abstract: Detailed image captioning demands both factual grounding and fine-grained coverage, yet existing methods have struggled to achieve them simultaneously.

Revisiting the Reliability of Language Models in Instruction-Following

Model ReleasesDGX agent

arXiv:2512.14754v2 Announce Type: replace-cross Abstract: Advanced LLMs have achieved near-ceiling instruction-following accuracy on benchmarks such as IFEval. However, these impressive scores do not

Robust Explanations for User Trust in Enterprise NLP Systems

Model ReleasesDGX agent

arXiv:2604.12069v1 Announce Type: cross Abstract: Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common case of black

Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss

Model ReleasesDGX agent

arXiv:2604.12911v1 Announce Type: cross Abstract: Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to p

RPG-SAM: Reliability-Weighted Prototypes and Geometric Adaptive Threshold Selection for Training-Free One-Shot Polyp Segmentation

Model ReleasesDGX agent

arXiv:2603.07436v2 Announce Type: replace Abstract: Training-free one-shot segmentation offers a scalable alternative to expert annotations where knowledge is often transferred from support images and

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

Model ReleasesDGX agent

arXiv:2509.18127v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) enable interpretability research by decomposing entangled model activations into monosemantic features. However, un

Scale-aware Message Passing For Graph Node Classification

Model ReleasesDGX agent

arXiv:2411.19392v3 Announce Type: replace Abstract: Most Graph Neural Networks (GNNs) operate at the first-order scale, even though multi-scale representations are known to be crucial in domains such

SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker

Model ReleasesDGX agent

arXiv:2604.12502v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) in multimodal tracking reveals a concerning trend where recent performance gains are often achieved at the cost

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents

Model ReleasesDGX agent

arXiv:2510.10073v2 Announce Type: replace-cross Abstract: Large vision-language model (LVLM)-based web agents are emerging as powerful tools for automating complex online tasks. However, when deployed

See, Point, Refine: Multi-Turn Approach to GUI Grounding with Visual Feedback

Model ReleasesDGX agent

arXiv:2604.13019v1 Announce Type: new Abstract: Computer Use Agents (CUAs) fundamentally rely on graphical user interface (GUI) grounding to translate language instructions into executable screen acti

SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From

Model ReleasesDGX agent

arXiv:2509.26404v2 Announce Type: replace-cross Abstract: Fingerprinting Large Language Models (LLMs)is essential for provenance verification and model attribution. Existing fingerprinting methods are

Self-Adversarial One Step Generation via Condition Shifting

Model ReleasesDGX agent

arXiv:2604.12322v1 Announce Type: new Abstract: The push for efficient text to image synthesis has moved the field toward one step sampling, yet existing methods still face a three way tradeoff among

Self-Monitoring Benefits from Structural Integration: Lessons from Metacognition in Continuous-Time Multi-Timescale Agents

Model ReleasesDGX agent

arXiv:2604.11914v1 Announce Type: new Abstract: Self-monitoring capabilities -- metacognition, self-prediction, and subjective duration -- are often proposed as useful additions to reinforcement learn

SEW: Self-Evolving Agentic Workflows for Automated Code Generation

Model ReleasesDGX agent

arXiv:2505.18646v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated effectiveness in code generation tasks. To enable LLMs to address more complex coding challenge

Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2603.01045v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed in multi-agent systems to overcome context limitations by distributing information across agen

SinkSAM-Net: Knowledge-Driven Self-Supervised Sinkhole Segmentation Using Topographic Priors and Segment Anything Model

Model ReleasesDGX agent

arXiv:2410.01473v2 Announce Type: replace Abstract: Soil sinkholes significantly influence soil degradation, infrastructure vulnerability, and landscape evolution. However, their irregular shapes, com

SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents

Model ReleasesDGX agent

arXiv:2604.12040v1 Announce Type: cross Abstract: We present SIR-Bench, a benchmark of 794 test cases for evaluating autonomous security incident response agents that distinguishes genuine forensic in

SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks

Model ReleasesDGX agent

arXiv:2506.14512v4 Announce Type: replace Abstract: Large Language Models (LLMs) have undergone rapid progress, largely attributed to reinforcement learning on complex reasoning tasks. In contrast, wh

Socrates Loss: Unifying Confidence Calibration and Classification by Leveraging the Unknown

Model ReleasesDGX agent

arXiv:2604.12245v1 Announce Type: cross Abstract: Deep neural networks, despite their high accuracy, often exhibit poor confidence calibration, limiting their reliability in high-stakes applications.

Space Force looks at moving 'significant number' of launches from ULA to SpaceX

Model ReleasesDGX agent

The US Space Force is reassessing its launch strategy following repeated technical malfunctions experienced by ULA's Vulcan rocket, particularly issues with its solid boosters. These reliability conce

SpaceX is moving at literal light speed right now SpaceX launched 2 more Falcon 9s in under 24 hours AGAIN - Florida in the morning, Califor…

Model ReleasesDGX agent

SpaceX is moving at literal light speed right now SpaceX launched 2 more Falcon 9s in under 24 hours AGAIN - Florida in the morning, California at night Even with unlimited budgets, government space p

Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping

Model ReleasesDGX agent

arXiv:2603.23998v2 Announce Type: replace Abstract: Existing approaches to increasing the effective depth of Transformers predominantly rely on parameter reuse, extending computation through recursive

SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model

Model ReleasesDGX agent

arXiv:2511.22039v3 Announce Type: replace Abstract: This paper introduces a novel architecture for trajectory-conditioned forecasting of future 3D scene occupancy. In contrast to methods that rely on

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

Model ReleasesDGX agent

arXiv:2604.12102v1 Announce Type: new Abstract: We introduce compute-grounded reasoning (CGR), a design paradigm for spatial-aware research agents in which every answerable sub-problem is resolved by

Spotify launches a feature to buy physical books in the US and UK, powered by Bookshop.org, and expands its Page Match tool to support 30 additional languages (Lauren Forristal/TechCrunch)

Model ReleasesDGX agent

Lauren Forristal / TechCrunch: Spotify launches a feature to buy physical books in the US and UK, powered by Bookshop.org, and expands its Page Match tool to support 30 additional languages — In Febru

.@Starlink is the best internet around. On flights it is miles ahead of anything else. Just a superior service/product.

Model ReleasesDGX agent

Starlink, SpaceX's satellite internet service, has been praised for delivering superior connectivity compared to competing options, including in-flight internet services on commercial aircraft. The se

StoryScope: Investigating idiosyncrasies in AI fiction

Model ReleasesDGX agent

arXiv:2604.03136v4 Announce Type: replace Abstract: As AI-generated fiction becomes increasingly prevalent, questions of authorship and originality are becoming central to how written work is evaluate

Style-Decoupled Adaptive Routing Network for Underwater Image Enhancement

Model ReleasesDGX agent

arXiv:2604.12257v1 Announce Type: new Abstract: Underwater Image Enhancement (UIE) is essential for robust visual perception in marine applications. However, existing methods predominantly rely on uni

Subspace-Guided Feature Reconstruction for Unsupervised Anomaly Localization

Model ReleasesDGX agent

arXiv:2309.13904v3 Announce Type: replace Abstract: Unsupervised anomaly localization aims to identify anomalous regions that deviate from normal sample patterns. Most recent methods perform feature m

SynthPix: A lightspeed PIV image generator

Model ReleasesDGX agent

arXiv:2512.09664v2 Announce Type: replace-cross Abstract: We describe SynthPix, a synthetic image generator for Particle Image Velocimetry (PIV) with a focus on performance and parallelism on accelera

T2I-BiasBench: A Multi-Metric Framework for Auditing Demographic and Cultural Bias in Text-to-Image Models

Model ReleasesDGX agent

arXiv:2604.12481v1 Announce Type: new Abstract: Text-to-image (T2I) generative models achieve impressive visual fidelity but inherit and amplify demographic imbalances and cultural biases embedded in

TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning

Model ReleasesDGX agent

arXiv:2604.12891v1 Announce Type: new Abstract: Deep learning (DL) compilers rely on cost models and auto-tuning to optimize tensor programs for target hardware. However, existing approaches depend on

TeRA: Vector-based Random Tensor Network for High-Rank Adaptation of Large Language Models

Model ReleasesDGX agent

arXiv:2509.03234v2 Announce Type: replace Abstract: Parameter-Efficient Fine-Tuning (PEFT) methods, such as Low-Rank Adaptation (LoRA), have significantly reduced the number of trainable parameters ne

Tesla has just released a new app update. Clips downloaded to your phone now finally include details such as speed, steering wheel angle, an…

Model ReleasesDGX agent

Tesla has just released a new app update. Clips downloaded to your phone now finally include details such as speed, steering wheel angle, and Self-Driving state. No more screen recording FSD clips! He

The example prompt for Google's new Gemini Flash TTS text-to-speed model is a lot https://simonwillison.net/2026/Apr/15/gemini-31-flash-tts/

Model ReleasesDGX agent

Google's Gemini 3.1 Flash TTS (text-to-speech) model includes a notably elaborate or extensive example prompt, which Simon Willison highlighted as noteworthy. The post likely comments on the complexit

The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break

Model ReleasesDGX agent

arXiv:2604.11978v1 Announce Type: new Abstract: Large language model (LLM) agents perform strongly on short- and mid-horizon tasks, but often break down on long-horizon tasks that require extended, in

The next evolution of the Agents SDK

Model ReleasesDGX agent

OpenAI's Agents SDK represents an evolution of their framework for building agentic AI applications, providing developers with tools to create, orchestrate, and deploy AI agents that can perform multi

The Verification Tax: Fundamental Limits of AI Auditing in the Rare-Error Regime

Model ReleasesDGX agent

arXiv:2604.12951v1 Announce Type: new Abstract: The most cited calibration result in deep learning -- post-temperature-scaling ECE of 0.012 on CIFAR-100 (Guo et al., 2017) -- is below the statistical

This is becoming a pattern in AI that makes talking about capabilities challenging. First, there are overstated claims (like the flubbed Erd…

Model ReleasesDGX agent

This is becoming a pattern in AI that makes talking about capabilities challenging. First, there are overstated claims (like the flubbed Erdos problems last year), then minor wins (AI helps with disco

Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems

Model ReleasesDGX agent

arXiv:2604.12231v1 Announce Type: new Abstract: Large language models (LLMs) have transformed AI research thanks to their powerful internal capabilities and knowledge. However, existing LLMs still fai

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Model ReleasesDGX agent

arXiv:2604.12012v1 Announce Type: new Abstract: Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classificat

← Previous
1…347348349350351…372
Next →