AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

86,457Total entries
1Added by human
86,456Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,039 results
7 Jul 2026

Diffusion learning reveals viable parameter manifolds and compensation geometry in biological dynamical systems

Model ReleasesDGX agent

arXiv:2607.03671v1 Announce Type: cross Abstract: Models of complex systems often have many parameters, yet are constrained by far fewer experimentally accessible observables: similar activity can eme

Double Fuzzy Probabilistic Interval Linguistic Term Set and a Dynamic Fuzzy Decision Making Model based on Markov Process with tts Application in Multiple Criteria Group Decision Making

ResearchDGX agent

arXiv:2111.15255v2 Announce Type: replace-cross Abstract: The probabilistic linguistic term has been proposed to deal with probability distributions in provided linguistic evaluations. However, becaus

Effective Distillation to Hybrid xLSTM Architectures

Model Releases
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2603.15590v2 Announce Type: replace Abstract: There have been numerous attempts to distill quadratic attention-based large language models (LLMs) into sub-quadratic linearized architectures. How

Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising

Model ReleasesDGX agent

arXiv:2607.04653v1 Announce Type: new Abstract: While modern video diffusion models excel in visual fidelity, maintaining long-range physical consistency remains a formidable challenge. Conventional p

IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning

Model ReleasesDGX agent

arXiv:2607.03614v1 Announce Type: new Abstract: Spatial question answering is the dominant paradigm for evaluating spatial intelligence in Vision-Language Models (VLMs), but it leaves a complementary

Inelastic Constitutive Kolmogorov-Arnold Networks: A generalized framework for automated discovery of interpretable inelastic material models

ResearchDGX agent

arXiv:2602.17750v2 Announce Type: replace-cross Abstract: A key problem of solid mechanics is the identification of the constitutive law of a material, that is, the relation between strain history and

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation

Model ReleasesDGX agent

arXiv:2512.10730v2 Announce Type: replace Abstract: Recent advances in motion-aware large language models have shown remarkable promise for jointly learning motion understanding and generation knowled

Language Models as Higher-Order Planning Formalizers

ResearchDGX agent

arXiv:2603.23844v2 Announce Type: replace Abstract: Recent work provides overwhelming evidence that LLMs, even those trained to scale their reasoning trace, quickly deteriorate at planning as problems

LLMs Encode Harmfulness and Refusal Separately

Model ReleasesDGX agent

arXiv:2507.11878v5 Announce Type: replace Abstract: LLMs are trained to refuse harmful instructions, but do they truly understand harmfulness beyond just refusing? Prior work has shown that LLMs' refu

miMamba: EEG-based Emotion Recognition with Multi-scale Inverted Mamba Models

ResearchDGX agent

arXiv:2409.07589v2 Announce Type: cross Abstract: EEG-based emotion recognition holds significant potential in the field of brain-computer interfaces. A key challenge lies in extracting discriminative

Model Confidence-Guided Multi-Image Fusion of Fundus Images for Diabetic Retinopathy Diagnosis

ResearchDGX agent

arXiv:2607.03643v1 Announce Type: cross Abstract: Purpose: Early screening for eye diseases is critical in low- and middle-income countries where access to care is limited. We investigate whether a co

OpenSIR: Open-Ended Self-Improving Reasoner

ResearchDGX agent

arXiv:2511.00602v4 Announce Type: replace Abstract: Recent advances in large language model (LLM) reasoning through reinforcement learning rely on annotated datasets for verifiable rewards, which may

Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters

Model ReleasesDGX agent

arXiv:2504.08791v3 Announce Type: replace-cross Abstract: On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low thr

Sample-Efficient Pareto Front Modeling for Energy-Aware Reinforcement Learning Using Bayesian Optimization

SafetyDGX agent

arXiv:2607.03140v1 Announce Type: new Abstract: Industrial automation increasingly demands control strategies that balance operational performance with strict energy efficiency requirements. A common

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning

Model ReleasesDGX agent

arXiv:2603.23483v2 Announce Type: replace-cross Abstract: Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through

The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs

Model ReleasesDGX agent

arXiv:2607.02587v1 Announce Type: cross Abstract: Model cards quote trust-benchmark scores without recording when they were measured, and the same number is routinely carried across successive checkpo

Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation

Model ReleasesDGX agent

arXiv:2603.17673v2 Announce Type: replace-cross Abstract: LLM agents are becoming increasingly important in the security domain, but leading systems are often closed-source, cloud-based, hard to repro

TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-Checking

Model ReleasesDGX agent

arXiv:2410.15135v5 Announce Type: replace Abstract: With the surge of online misinformation, Large Language Models (LLMs) and Reasoning Large Language Models (RLMs) serving as Automatic Fact-Checking

UniVideo: Unified Understanding, Generation, and Editing for Videos

Model ReleasesDGX agent

arXiv:2510.08377v4 Announce Type: replace Abstract: Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain.

4 Jul 2026

Only AI can keep up with AI. That's what will guide the entire security model of the next decade.

TutorialsDGX agent

AI-powered security systems will be necessary to detect and defend against AI-based threats, as human security experts cannot match the speed and sophistication of AI attacks. This perspective suggest

spending the last week at @aidotengineer was awesome. too many great convos to cover them all, but jotted down some things that stood out: -…

Model ReleasesDGX agent

spending the last week at @aidotengineer was awesome. too many great convos to cover them all, but jotted down some things that stood out: - Lots of discussion around open source models. I spoke with

3 Jul 2026

An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms

Model ReleasesDGX agent

arXiv:2603.29466v2 Announce Type: replace-cross Abstract: Existing methods for quantifying predictive uncertainty in neural networks are either computationally intractable for large language models or

Distributionally Robust Listwise Preference Optimization

Model ReleasesDGX agent

arXiv:2607.01715v1 Announce Type: new Abstract: Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, o

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

Model ReleasesDGX agent

arXiv:2607.01751v1 Announce Type: cross Abstract: Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right ti

Meta-Benchmarks for Financial-Services LLM Evaluation

Model ReleasesDGX agent

arXiv:2607.01740v1 Announce Type: new Abstract: Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services work: a model th

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

Model ReleasesDGX agent

arXiv:2607.02047v1 Announce Type: cross Abstract: Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts.

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling

Model ReleasesDGX agent

arXiv:2607.01612v1 Announce Type: new Abstract: Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering

SCAPE: Accurate and Efficient LLM Training with Extreme Sparse Communication

Model ReleasesDGX agent

arXiv:2607.01678v1 Announce Type: new Abstract: Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, w

2 Jul 2026

Continuous Speculative Decoding for Autoregressive Image Generation

SafetyDGX agent

arXiv:2411.11925v3 Announce Type: replace Abstract: Continuous visual autoregressive (AR) models have demonstrated promising performance in image generation, but their inherently sequential nature res

DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning

Model ReleasesDGX agent

arXiv:2607.00341v1 Announce Type: cross Abstract: Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought (CoT). How

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

Model ReleasesDGX agent

arXiv:2512.15702v2 Announce Type: replace Abstract: Autoregressive video diffusion models hold promise for world simulation but are vulnerable to exposure bias arising from the train-test mismatch. Wh

From 'Strings' to 'Things' for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems

Model ReleasesDGX agent

arXiv:2607.00003v1 Announce Type: cross Abstract: Personal Knowledge Graphs (PKGs) offer a privacy-preserving framework for modeling user preferences, yet constructing them from unstructured, decentra

LLM-Guided ODE Discovery and Parameter Inference from Small-Cohort Aggregate Data

Model ReleasesDGX agent

arXiv:2607.00733v1 Announce Type: cross Abstract: Mechanistic modeling via ordinary differential equations (ODEs) provides interpretable descriptions of complex dynamics and enables inference of under

MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos

Model ReleasesDGX agent

arXiv:2607.00491v1 Announce Type: cross Abstract: Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Exis

OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection

Model ReleasesDGX agent

arXiv:2505.19889v3 Announce Type: replace Abstract: Visual fall detection models are usually trained on small, staged datasets. Their real-world utility remains unclear; such data lacks diversity and

PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

Model ReleasesDGX agent

arXiv:2607.00436v1 Announce Type: new Abstract: Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more

Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps

Model ReleasesDGX agent

arXiv:2607.00004v1 Announce Type: cross Abstract: While advanced foundation models like ModernBERT significantly outperform older architectures in dense retrieval, they surprisingly lag behind the agi

1 Jul 2026

Amplifying Membership Signal Through Chained Regeneration

ResearchDGX agent

arXiv:2606.31991v1 Announce Type: cross Abstract: The tendency of large generative models to memorize training data makes sample verification critical for privacy auditing and copyright enforcement. C

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

Model ReleasesDGX agent

arXiv:2606.30850v1 Announce Type: new Abstract: Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic unce

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

Model ReleasesDGX agent

arXiv:2606.22723v2 Announce Type: replace Abstract: Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended,

Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling

Local AiDGX agent

arXiv:2606.31844v1 Announce Type: cross Abstract: A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in globally observab

Claude Fable 5 will be available again globally tomorrow. After a series of productive conversations with the US government, we're redeployi…

Model ReleasesDGX agent

Claude Fable 5 will be available again globally tomorrow. After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and blo

Content Independence Day, one year on: building the business model for the agentic Internet

AgentsDGX agent

One year after declaring Content Independence Day, a dynamic market for monetized content has officially emerged. In this report, we examine how the rise of autonomous AI agents is upending traditiona

CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM

ApplicationsDGX agent

arXiv:2606.31332v1 Announce Type: new Abstract: Protein automodeling from cryo-EM density maps faces unique challenges in enforcing physicochemical validity and managing conformational heterogeneity.

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning

ApplicationsDGX agent

arXiv:2606.31257v1 Announce Type: new Abstract: The standard way to read latent knowledge out of a model, a linear probe confirmed by a steering recovery, can systematically overstate what a vision-la

ElemeNet: Multiscale Molecular Machine Learning with Uncertainty Quantification Across the Periodic Table

Model ReleasesDGX agent

arXiv:2606.30961v1 Announce Type: cross Abstract: Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models ha

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking

SafetyDGX agent

arXiv:2509.12046v2 Announce Type: replace-cross Abstract: Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned gen

Mind the Residual Gap: Probabilistic Downscaling under Real-World Bias

Model ReleasesDGX agent

arXiv:2606.30821v1 Announce Type: new Abstract: Probabilistic downscaling is the task of modeling the conditional distribution of high-resolution fields given coarse inputs, and is a central challenge

Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving

AgentsDGX agent

arXiv:2606.31160v1 Announce Type: new Abstract: Modern Vision-Language-Action (VLA) planners for autonomous driving emit a chain-of-causation (CoC) reasoning step before producing a trajectory. The re

Safely Releasing Frontier Models to Customers

IndustryDGX agent

It’s our goal for AWS to be the most secure place to run any workload, and in support of that we’ve been deeply investing in security across our services since AWS's inception more than two decades ag

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA

Local AiDGX agent

arXiv:2606.32002v1 Announce Type: new Abstract: Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions about a document, answers them fro

Stage-Transition Dense Reward Modeling for Reinforcement Learning

ResearchDGX agent

arXiv:2606.31377v1 Announce Type: cross Abstract: Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping si

Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies

Model ReleasesDGX agent

arXiv:2606.31039v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies re

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

ResearchDGX agent

arXiv:2606.31128v1 Announce Type: cross Abstract: Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-lev

VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

SafetyDGX agent

arXiv:2603.16271v3 Announce Type: replace Abstract: Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial d

30 Jun 2026

Agent-Computer Observation Interfaces Enable Dynamic Computer Use

Model ReleasesDGX agent

arXiv:2606.29472v1 Announce Type: new Abstract: SWE-agent established the action interface as an underexplored design axis for software-engineering agents; we make the analogous case for the observati

BERTomelo: Your Portuguese Encoder Best Friend

Model ReleasesDGX agent

arXiv:2606.28999v1 Announce Type: cross Abstract: Encoders have become the state of the art for multiple NLP tasks, especially those requiring deep contextual understanding. While multilingual models

Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation

Model ReleasesDGX agent

arXiv:2511.05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning remain

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

Model ReleasesDGX agent

arXiv:2606.29920v1 Announce Type: new Abstract: Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliab

Can LLMs Reliably Self-Report Adversarial Prefills, and How?

SafetyDGX agent

arXiv:2606.23671v2 Announce Type: replace Abstract: Prior work shows that large language models (LLMs) exhibit introspective capability on benign tasks. We extend the question to safety contexts and e

← Previous
1…275276277278279…1034
Next →