AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
12 May 2026

EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation

Model ReleasesDGX agent

arXiv:2605.09378v1 Announce Type: cross Abstract: Long-horizon video generation has advanced in visual quality, yet existing methods still struggle to maintain knowledge consistency and coherent pedag

Efficient Ensemble Selection from Binary and Pairwise Feedback

Model ReleasesDGX agent

arXiv:2605.09588v1 Announce Type: cross Abstract: Organizations increasingly deploy multiple AI systems across task domains, but selecting a small, high-performing ensemble can require costly model ca

Efficient Evaluation of LLM Performance with Statistical Guarantees

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2601.20251v3 Announce Type: replace-cross Abstract: Exhaustively evaluating many large language models (LLMs) on a large suite of benchmarks is expensive. We cast benchmarking as finite-populati

Efficient Neural Architectures for Real-Time ECG Interpretation on Limited Hardware

Model ReleasesDGX agent

arXiv:2605.09848v1 Announce Type: new Abstract: Electrocardiogram (ECG) interpretation is essential for diagnosing a wide range of cardiac abnormalities. While deep learning has shown strong potential

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

Model ReleasesDGX agent

arXiv:2605.09874v1 Announce Type: cross Abstract: Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more

EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents

Model ReleasesDGX agent

arXiv:2605.10332v1 Announce Type: new Abstract: Embodied agents can benefit from skills that guide object search, action execution, and state changes across diverse environments. Since embodied enviro

EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding

Model ReleasesDGX agent

arXiv:2605.08847v1 Announce Type: new Abstract: In the context of today's high-pressure, aging society, the demand for large-scale emotional models capable of providing empathetic support is more crit

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.09826v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to track others epistemic state, makes humans efficient collaborators. AI agents need the same capacity in multi agent

EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving

Model ReleasesDGX agent

arXiv:2605.10556v1 Announce Type: new Abstract: As large language models span dense, mixture-of-experts, and state-space architectures and are deployed on heterogeneous accelerators under increasingly

EpiGraph: A Knowledge Graph and Benchmark for Evidence-Intensive Reasoning in Epilepsy

Model ReleasesDGX agent

arXiv:2605.09505v1 Announce Type: new Abstract: Epilepsy diagnosis and treatment require evidence-intensive reasoning across heterogeneous clinical knowledge, including biosignal patterns, genetic mec

Equilibrium Residuals Expose Three Regimes of Matrix-Game Strategic Reasoning in Language Models

Model ReleasesDGX agent

arXiv:2605.10410v1 Announce Type: new Abstract: Large language models can score well on named game-theory benchmarks while failing on the same strategic computation once semantic cues are removed. We

ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

Model ReleasesDGX agent

arXiv:2505.22919v3 Announce Type: replace Abstract: Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of 'clinical re

ERIS: Enhancing Privacy and Scalability in Federated Learning via Federated Shard Aggregation

Model ReleasesDGX agent

arXiv:2602.08617v2 Announce Type: replace Abstract: Scaling Federated Learning (FL) to billion-parameter models forces a challenging trade-off between privacy, scalability, and model utility. Existing

Even @haider1 sees that Mythos has been overhyped.

Model ReleasesDGX agent

Even @haider1 sees that Mythos has been overhyped. mythos is pretty on par with gpt-5.5 and while gpt-5.5 is currently SOTA, it's not anything like what anthropic describes mythos as it's pretty obvio

EverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA

Model ReleasesDGX agent

arXiv:2510.06371v2 Announce Type: replace-cross Abstract: Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries re

Exactness Matters for Physical Rule Enforcement

Model ReleasesDGX agent

arXiv:2605.08285v1 Announce Type: new Abstract: Autoregressive scientific forecasters often enforce physical or structural constraints by repairing each predicted state before feeding it back into the

Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups

Model ReleasesDGX agent

arXiv:2605.08671v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed not only to make decisions but to explain them. While AI decision fairness has been studied ext

Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness

Model ReleasesDGX agent

arXiv:2509.13332v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly adopted as automated judges in benchmarking and reward modeling, ensuring their reliability, effici

Exploitation Without Deception: Dark Triad Feature Steering Reveals Separable Antisocial Circuits in Language Models

Model ReleasesDGX agent

arXiv:2605.09773v1 Announce Type: cross Abstract: We use sparse autoencoder (SAE) feature steering to amplify Dark Triad personality traits (Machiavellianism, narcissism, and psychopathy) in Llama-3.3

Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?

Model ReleasesDGX agent

arXiv:2603.00166v2 Announce Type: replace-cross Abstract: Recent advances in generative AI have shown human-level performance in complex content creation. However, we identify a 'Paradox of Simplicity

expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling

Model ReleasesDGX agent

arXiv:2605.09923v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, where Group Relative Policy Optim

FactoryNet: A Large-Scale Dataset toward Industrial Time-Series Foundation Models

Model ReleasesDGX agent

arXiv:2605.09081v1 Announce Type: cross Abstract: We introduce the first universal pretraining corpus for industrial time-series data: FactoryNet. 51M datapoints across 23k end-to-end task executions

Failing Forward: Adaptive Failure-Informed Learning for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.08434v1 Announce Type: new Abstract: Vision-language-action (VLA) models provide a promising paradigm for scalable robotic manipulation, yet their reliance on success-only behavioral clonin

Falcon 9 launches NROL-172 to orbit from pad 4E in California

Model ReleasesDGX agent

SpaceX's Falcon 9 rocket successfully launched the NROL-172 classified national reconnaissance payload from Space Launch Complex 4E at Vandenberg Space Force Base in California. This mission was condu

Fashion Florence: Fine-Tuning Florence-2 for Structured Fashion Attribute Extraction

Model ReleasesDGX agent

arXiv:2605.09827v1 Announce Type: cross Abstract: We present Fashion Florence, a Florence-2 vision-language model fine-tuned with LoRA to extract structured fashion attributes from clothing images. Gi

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition

Model ReleasesDGX agent

arXiv:2605.10127v1 Announce Type: new Abstract: Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image

Fast mode for Claude Opus 4.7 is now available in Cursor! It's 2.5x the speed at 6x the cost. For most tasks, we recommend using the standar…

Model ReleasesDGX agent

Cursor has released fast mode for Claude Opus 4.7, offering 2.5x faster processing speeds but at 6x the cost compared to standard mode. The announcement suggests that for most tasks, standard mode rem

Fast mode for Claude Opus 4.7 is now available in research preview on the API and in Claude Code.

Model ReleasesDGX agent

Anthropic has released a fast mode for Claude Opus 4.7, now available in research preview for both the API and Claude Code, offering improved performance for compatible workloads. This feature allows

Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking

Model ReleasesDGX agent

arXiv:2605.08119v1 Announce Type: cross Abstract: Tian (2025) proves a repulsion theorem (Theorem 6) for the matrix B = (widetilde{F}^op widetilde{F} + eta I)^{-1} during the interactive feature-learn

Feature Rivalry in Sparse Autoencoder Representations: A Mechanistic Study of Uncertainty-Driven Feature Competition in LLMs

Model ReleasesDGX agent

arXiv:2605.08149v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) decompose large language model representations into interpretable features, but how these features interact under uncertain

Featurized Occupation Measures for Structured Global Search in Numerical Optimal Control

Model ReleasesDGX agent

arXiv:2603.16231v2 Announce Type: replace-cross Abstract: Numerical optimal control has long been split between globally structured but dimensionally intractable Hamilton--Jacobi--Bellman (HJB) method

Federated Language Models Under Bandwidth Budgets: Distillation Rates and Conformal Coverage

Model ReleasesDGX agent

arXiv:2605.09986v1 Announce Type: cross Abstract: Training a language model on data scattered across bandwidth-limited nodes that cannot be centralized is a setting that arises in clinical networks, e

Filtering Memorization from Parameter-Space in Diffusion Models

Model ReleasesDGX agent

arXiv:2605.10439v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become a widely used mechanism for customizing diffusion models, enabling users to inject new visual concepts or styles t

Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain

Model ReleasesDGX agent

arXiv:2605.09106v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in financial contexts, raising critical concerns about reliability, alignment, and susceptibility

FinTSB: A Comprehensive and Practical Benchmark for Financial Time Series Forecasting

Model ReleasesDGX agent

arXiv:2502.18834v2 Announce Type: replace-cross Abstract: Financial time series (FinTS) record the behavior of human-brain-augmented decision-making, capturing valuable historical information that can

Fitting Multilinear Polynomials for Logic Gate Networks

Model ReleasesDGX agent

arXiv:2605.08657v1 Announce Type: cross Abstract: We study learnable logic gate networks that stack layers of 2-input Boolean gates to build combinational circuits. Every 2-input gate has a unique mul

Fix the Loss, Not the Radius: Rethinking the Adversarial Perturbation of Sharpness-Aware Minimization

Model ReleasesDGX agent

arXiv:2605.10183v1 Announce Type: new Abstract: Sharpness-Aware Minimization (SAM) improves generalization by minimizing the worst-case loss within a fixed parameter-space radius neighborhood. SAM and

FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

Model ReleasesDGX agent

arXiv:2605.09355v1 Announce Type: new Abstract: Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: (1) multi-task pretraining, t

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models

Model ReleasesDGX agent

arXiv:2605.09218v1 Announce Type: cross Abstract: 3D scene understanding spans reasoning about free space, object grounding, hypothetical object insertions, complex geometric relationships, and integr

FlashClear: Ultra-Fast Image Content Removal via Efficient Step Distillation and Feature Caching

Model ReleasesDGX agent

arXiv:2605.09003v1 Announce Type: new Abstract: Recently, diffusion-based object removal models have achieved impressive results in eliminating objects and their associated visual effects. However, th

Follow the Mean: Reference-Guided Flow Matching

Model ReleasesDGX agent

arXiv:2605.10302v1 Announce Type: new Abstract: Existing approaches to controllable generation typically rely on fine-tuning, auxiliary networks, or test-time search. We show that flow matching admits

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs

Model ReleasesDGX agent

arXiv:2605.08905v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success on reasoning benchmarks through Reinforcement Learning with Verifiable Rewards (RLVR), exc

FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models

Model ReleasesDGX agent

arXiv:2605.10141v1 Announce Type: new Abstract: Recent neural theorem provers use reinforcement learning with verifiable rewards (RLVR), where proof assistants provide binary correctness signals. Whil

FORTIS: Benchmarking Over-Privilege in Agent Skills

Model ReleasesDGX agent

arXiv:2605.09163v1 Announce Type: new Abstract: Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This

Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling

Model ReleasesDGX agent

arXiv:2512.02010v5 Announce Type: replace Abstract: As large language models have grown larger, interest has grown in low-precision numerical formats such as NVFP4 as a way to improve speed and reduce

FPGA-Based Hardware Architecture for Contrast Maximization in Event-Based Vision

Model ReleasesDGX agent

arXiv:2605.09581v1 Announce Type: new Abstract: This paper presents a hardware architecture that implements the Contrast Maximization (CM) algorithm in Field-Programmable Gate Array (FPGA) resources f

FRACTAL: SSM with Fractional Recurrent Architecture for Computational Temporal Analysis of Long Sequences

Model ReleasesDGX agent

arXiv:2605.08833v1 Announce Type: new Abstract: Effective sequence modeling fundamentally requires balancing the retention of unbounded history with the high-resolution detection of abrupt short-term

Frame In, Frame Out: Measuring Framing Bias in LLM-Generated News Summaries

Model ReleasesDGX agent

arXiv:2505.05406v2 Announce Type: replace Abstract: News headlines and summaries shape how events are interpreted through selective emphasis and omission, a phenomenon commonly referred to as framing.

FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence

Model ReleasesDGX agent

arXiv:2605.08820v1 Announce Type: cross Abstract: Artificial Intelligence (AI)-generated images have become increasingly realistic and readily adaptable to concrete real-world claims, creating new cha

FreeMOCA: Memory-Free Continual Learning for Malicious Code Analysis

Model ReleasesDGX agent

arXiv:2605.09664v1 Announce Type: cross Abstract: As over 200 million new malware samples are identified each year, antivirus systems must continuously adapt to the evolving threat landscape. However,

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation

Model ReleasesDGX agent

arXiv:2605.08712v1 Announce Type: new Abstract: Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimension

From Pixels to Concepts: Do Segmentation Models Understand What They Segment?

Model ReleasesDGX agent

arXiv:2605.09591v1 Announce Type: new Abstract: Segmentation is a fundamental vision task underlying numerous downstream applications. Recent promptable segmentation models, such as Segment Anything M

From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages

Model ReleasesDGX agent

arXiv:2605.09147v1 Announce Type: cross Abstract: Part-of-speech (POS) tagging for Medieval Romance languages remains challenging due to orthographic variation, morphological complexity, and limited a

Function-Space ADMM for Decentralized Federated Learning: A Control Theoretic Perspective

Model ReleasesDGX agent

arXiv:2605.09356v1 Announce Type: new Abstract: Decentralized federated learning (FL) is a promising approach for training machine learning models on sensor networks, Internet of Things (IoT) devices,

GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives

Model ReleasesDGX agent

arXiv:2605.09027v1 Announce Type: new Abstract: In multi-agent systems (MAS), a single deceptive agent can nullify all gains of an agentic AI collective and evade deployed defenses. However, existing

Gemini’s biggest new features are all about controlling your phone

Model ReleasesDGX agent

It is, once again, Gemini season. Google is announcing a host of new Gemini features during its pre-I/O Android showcase, many of which aim to help use your phone for you. You'll find Gemini in more p

General Agent Evaluation

Model ReleasesDGX agent

arXiv:2602.22953v2 Announce Type: replace Abstract: General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measur

Generating Leakage-Free Benchmarks for Robust RAG Evaluation

Model ReleasesDGX agent

arXiv:2605.08838v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is widely used to augment large language models (LLMs) with external knowledge. However, many benchmark datasets,

Generating Symmetric Materials using Latent Flow Matching

Model ReleasesDGX agent

arXiv:2605.10115v1 Announce Type: new Abstract: Tackling the task of materials generation, we aim to enhance the previously proposed All-atom Diffusion Transformer (ADiT) by introducing SymADiT, a sym

Generative Actor-Critic with Soft Bridge Policies

Model ReleasesDGX agent

arXiv:2605.08733v1 Announce Type: new Abstract: Expressive generative policies such as diffusion and flow models are appealing for MaxEnt online reinforcement learning because of their ability to mode

← Previous
1…262263264265266…377
Next →