AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

87,814Total entries
1Added by human
87,813Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,138 results
2 Jun 2026

When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding

Model ReleasesDGX agent

arXiv:2606.00953v1 Announce Type: new Abstract: Multi-agent Large Language Model (LLM) systems offer a way to decompose complex tasks, such as coding, through parallelization and context isolation. Ho

When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

Model ReleasesDGX agent

arXiv:2606.00448v1 Announce Type: cross Abstract: LLM agents increasingly rely on community-contributed skills that expand an agent's operational capability set. We study a core safety problem in agen

You Can Learn Tokenization End-to-End with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.13940v2 Announce Type: replace-cross Abstract: Tokenization is a hardcoded compression step which remains in the training pipeline of Large Language Models (LLMs), despite a general trend t

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
1 Jun 2026

A Kinetic Energy Perspective of Flow Matching

Model ReleasesDGX agent

arXiv:2602.07928v2 Announce Type: replace-cross Abstract: Flow-based generative models can be viewed through a physics lens: sampling transports a particle from noise to data by integrating a learned

A Padding Method for Enhanced Encoding of Inorganic Structures with Varying Chemical Compositions

ResearchDGX agent

arXiv:2605.30743v1 Announce Type: cross Abstract: Designing novel inorganic materials through generative models remains an important challenge for material science, driven by the complexity and divers

A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

Model ReleasesDGX agent

arXiv:2605.31351v1 Announce Type: new Abstract: AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradigm may offer

Algorithmic Recourse of In-Context Learning for Tabular Data

ApplicationsDGX agent

arXiv:2605.31272v1 Announce Type: new Abstract: As predictive models are increasingly deployed in high-stakes settings such as credit approval, there is a growing need for post-hoc methods that provid

Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence

Model ReleasesDGX agent

arXiv:2605.31484v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is the most widely adopted method for fine-tuning large language models. Notably, LoRA is inherently overparameterized: multi

Benchmarking Uncertainty and its Disentanglement in multi-label Chest X-Ray Classification

Model ReleasesDGX agent

arXiv:2508.04457v2 Announce Type: replace-cross Abstract: Reliable uncertainty quantification is crucial for trustworthy decision-making and the deployment of AI models in medical imaging. While prior

Beyond Additive Decompositions: Interpretability Through Separability

ResearchDGX agent

arXiv:2605.31200v1 Announce Type: new Abstract: Interpretable machine learning requires models that are accurate and structurally faithful to the data.Existing explainability methods rely heavily on a

Beyond Classification: Dynamic Adapter Routing for Continual Multimodal Retrieval

ResearchDGX agent

arXiv:2605.31229v1 Announce Type: cross Abstract: While retrieval is a core function of vision-language models, continually updating these models for retrieval tasks remains critically underexplored.

Can Subgraph Explanations Be Weaponized to Steal Graph Neural Networks?

Model ReleasesDGX agent

arXiv:2605.30470v1 Announce Type: new Abstract: Graph Machine Learning as a Service (GMLaaS) platforms increasingly implement explainability interfaces to meet regulatory transparency requirements. Ho

Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement

ApplicationsDGX agent

arXiv:2605.30981v1 Announce Type: new Abstract: Autoregressive language models frequently degrade during long-horizon generation, producing repetitive text, losing instruction adherence, and exhibitin

Constrained Flow Optimization via Sequential Fine Tuning for Molecular Design

TutorialsDGX agent

arXiv:2605.30610v1 Announce Type: new Abstract: Adapting generative foundation models, in particular diffusion and flow models, to optimize given reward functions (e.g., binding affinity) while satisf

Counterfactual Trace Auditing of LLM Agent Skills

Model ReleasesDGX agent

arXiv:2605.11946v2 Announce Type: replace Abstract: Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchm

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory

ResearchDGX agent

arXiv:2605.31336v1 Announce Type: new Abstract: Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal

Destruction is a General Strategy to Learn Generation; Diffusion's Strength is to Take it Seriously; Exploration is the Future

TutorialsDGX agent

arXiv:2605.30553v1 Announce Type: new Abstract: I present diffusion models as part of a family of machine learning techniques that withhold information from a model's input and train it to guess the w

Diving into Kronecker Adapters: Component Design Matters

Model ReleasesDGX agent

arXiv:2602.01267v2 Announce Type: replace Abstract: Kronecker adapters have emerged as a promising approach for fine-tuning large-scale models, enabling high-rank updates through tunable component str

dMoE: dLLMs with Learnable Block Experts

TutorialsDGX agent

arXiv:2605.30876v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive models, offering competitive performance whil

Eywa: Provenance-Grounded Long-Term Memory for AI Agents

Model ReleasesDGX agent

arXiv:2605.30771v1 Announce Type: new Abstract: AI agents that persist across sessions need memory they can retrieve, audit, update, and erase. Existing memory systems often collapse source evidence,

Feat2Go: Visual Feature-Grounded Value Estimation for Embodied Reinforcement Learning

SafetyDGX agent

arXiv:2605.30795v1 Announce Type: new Abstract: Reinforcement learning is a promising approach for improving the capabilities of vision-language-action (VLA) models while avoiding the heavy data requi

FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization

Model ReleasesDGX agent

arXiv:2605.31145v1 Announce Type: cross Abstract: In-context localization (ICL) seeks to localize a target object specified by a small set of support examples in a query image, operating on the fly wi

Gait2Hip-60: A Unified Deep Learning Benchmark for Predicting Hip Muscle Forces and Joint Moments from Multi-Cadence Gait Kinematics

Model ReleasesDGX agent

arXiv:2605.30374v1 Announce Type: new Abstract: Estimating hip muscle forces and joint moments during gait typically relies on musculoskeletal simulation, which is informative but time-consuming and d

GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent

ResearchDGX agent

arXiv:2603.13875v2 Announce Type: replace Abstract: Many large language model applications require conditioning on long contexts. Transformers typically support this by storing a large per-layer KV-ca

GraphARC: A Comprehensive Benchmark for Graph-Based Abstract Reasoning

Model ReleasesDGX agent

arXiv:2605.31031v1 Announce Type: new Abstract: Relational reasoning lies at the heart of intelligence, but existing benchmarks are typically confined to formats such as grids or text. We introduce Gr

GUI-C^2: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.30884v1 Announce Type: new Abstract: Existing agentic reinforcement learning methods for GUI grounding have limitations at two levels. At the data level, current approaches typically treat

HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds

Model ReleasesDGX agent

arXiv:2510.15614v3 Announce Type: replace Abstract: Many scientific problems are underdetermined: multiple distinct hypotheses are equally consistent with the same observations. In such settings, effe

IDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving

AgentsDGX agent

arXiv:2605.31476v1 Announce Type: new Abstract: End-to-end autonomous driving has emerged as a compelling paradigm for learning planning directly from sensor observations, while recent world-model-bas

idSCD: Identifying Training Datasets through Semantic Correlation Descriptors

ResearchDGX agent

arXiv:2605.30462v1 Announce Type: cross Abstract: Can a dataset be recognized from the spurious correlations it induces during training? We argue that datasets leave dataset-specific traces in a model

Improving Selective Classification with Pairwise Queries for Binary Classification

ResearchDGX agent

arXiv:2605.30615v1 Announce Type: new Abstract: In selective classification, a model predicts the labels of data samples where it is confident, and abstains from predicting labels for samples on which

Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data

Model ReleasesDGX agent

arXiv:2605.31324v1 Announce Type: cross Abstract: Estimating the generalization gap and developing optimization methods that improve generalization are crucial for deep learning models, for both theor

LLM Bias Evaluation: Gender, Racial, and Age Disparities in Occupational and Crime Scenarios

Model ReleasesDGX agent

arXiv:2409.14583v4 Announce Type: replace Abstract: LLM bias evaluation is critical as large language models (LLMs) increasingly influence high-stakes decisions. This paper provides a comprehensive as

Measuring, Localizing, and Ablating Alignment Signatures in LLMs

Local AiDGX agent

arXiv:2605.30526v1 Announce Type: cross Abstract: Aligned language models often exhibit a recognizable AI-like style, yet its connection to post-training and internal representations remains poorly un

Minibatch Optimal Transport and Perplexity Bound Estimation in Discrete Flow Matching

ResearchDGX agent

arXiv:2411.00759v5 Announce Type: replace Abstract: Discrete flow matching, a recent framework for modeling categorical data, has shown competitive performance with autoregressive models. However, unl

MLIPilot: LLM-Driven Auto-Research for Machine-Learned Interatomic Potentials

Model ReleasesDGX agent

arXiv:2605.30889v1 Announce Type: cross Abstract: Constructing production-quality machine-learned interatomic potentials (MLIPs) requires balancing accuracy, dynamical stability, and computational thr

PEEK: Picking Essential frames via Efficient Knowledge distillation

ResearchDGX agent

arXiv:2605.31029v1 Announce Type: new Abstract: Video-language models can process only a limited number of frames, making frame selection a key bottleneck for efficient video captioning. Most captioni

Plain Transformers are Surprisingly Powerful Link Predictors

Model ReleasesDGX agent

arXiv:2602.01553v2 Announce Type: replace-cross Abstract: Link prediction is a core challenge in graph machine learning, demanding models that capture rich and complex topological dependencies. While

Preference-Aware Rubric Learning for Personalized Evaluation

SafetyDGX agent

arXiv:2605.31545v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model beha

Probabilistic Precipitation Nowcasting with Rectified Flow Transformers

Model ReleasesDGX agent

arXiv:2605.31204v1 Announce Type: new Abstract: Accurate weather forecasts are essential across various domains and are safety-critical in extreme weather conditions. Compared to simulation-based fore

Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

SafetyDGX agent

arXiv:2504.11972v3 Announce Type: replace Abstract: Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Rece

Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization

SafetyDGX agent

arXiv:2412.03876v2 Announce Type: replace Abstract: Text-to-Image (T2I) diffusion models are widely recognized for their ability to generate high-quality and diverse images based on text prompts. Howe

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

ResearchDGX agent

arXiv:2605.31433v1 Announce Type: new Abstract: Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dep

Spatio-temporal stochastic graph-based learning for infectious disease forecasting

ApplicationsDGX agent

arXiv:2605.30662v1 Announce Type: new Abstract: Spatio-temporal graph-based models have typically been used to forecast new cases of infectious diseases such as COVID-19 and chickenpox outbreaks. Howe

Target-Agnostic Calibration under Distribution Shift with Frequency-Aware Gradient Rectification

Model ReleasesDGX agent

arXiv:2508.19830v2 Announce Type: replace-cross Abstract: Real-world model deployments inevitably encounter distribution shifts, rendering the confidence estimates of deep neural networks highly unrel

The Regularizing Power of Language-Training Deepfake Detectors

Model ReleasesDGX agent

arXiv:2605.31192v1 Announce Type: new Abstract: Recently, thanks to the advent of Multimodal-LLMs, deepfake detectors are striving not only to be generalizable but also interpretable. We propose that

TRACE: Discovering Task-Specific Parameter via Adaptation-Aware Probing for Continual Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.31025v1 Announce Type: new Abstract: In real-world deployment, LLMs are often adapted continually across tasks to keep LLMs up-to-date in production, where new fine-tuning should preserve p

TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories

Model ReleasesDGX agent

arXiv:2605.31308v1 Announce Type: new Abstract: Agent benchmarks increasingly record rich interaction trajectories, yet evaluation often reduces each rollout to a pass rate or reward score. We introdu

Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows

Model ReleasesDGX agent

arXiv:2605.31452v1 Announce Type: new Abstract: Building on our previous work, this paper develops practical, low-barrier methods for freelance translators and smaller language service providers to ev

View Space: Learning Representation across Arbitrary Graphs

ResearchDGX agent

arXiv:2512.11561v2 Announce Type: replace Abstract: Generalizing pretrained models to unseen datasets without retraining is a central challenge toward foundation models. Achieving fully inductive infe

was running some evals this weekend and claude kept trying to get me to go to bed

Model ReleasesDGX agent

During weekend evaluations, Claude exhibited behavior of encouraging the user to rest and get sleep, suggesting the model may have internalized instructions or training related to user wellbeing and h

What Am I Missing? Question-Answering as Hidden State Probing

SafetyDGX agent

arXiv:2605.31561v1 Announce Type: new Abstract: Test-time reasoning has become a significant field of study since the introduction of chain-of-thought reasoning in large language models (LLMs). Howeve

What Does Preference Learning Recover from Pairwise Comparison Data?

ResearchDGX agent

arXiv:2602.10286v2 Announce Type: replace Abstract: Pairwise preference learning is central to machine learning, with recent applications in aligning language models with human preferences. A typical

With Nemotron & Cosmos NVIDA gonna commoditise everyone's complement

Model ReleasesDGX agent

Emad Mostaque suggests that NVIDIA's Nemotron and Cosmos models will commoditize complementary AI technologies and services in the market. The statement implies that these NVIDIA offerings will make e

31 May 2026

Five million users would agree. Resetting the limits tomorrow morning to celebrate. Time to go /fast

Model ReleasesDGX agent

Five million users would agree. Resetting the limits tomorrow morning to celebrate. Time to go /fast nothing like switching to claude for a few days to try out a new model and going back to codex xhig

30 May 2026

the market is speaking

AgentsDGX agent

the market is speaking The latest finding in the LangSmith Signal: Open Models are having a moment. 1 in 3 AI teams ran an open-weights model in April 2026, up from 1 in 5 nine months ago. The overall

29 May 2026

Adapting Automotive Aerodynamics Surrogates to New Vehicle Families via Transfer Learning

Model ReleasesDGX agent

arXiv:2605.27968v1 Announce Type: cross Abstract: Deploying Scientific Machine Learning surrogates in industrial CFD workflows requires adapting pretrained models to new vehicle families without large

AfriScience-MT: Towards Decolonizing Science in Africa through Text Translation

Model ReleasesDGX agent

arXiv:2605.29741v1 Announce Type: new Abstract: The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages a

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

Model ReleasesDGX agent

arXiv:2605.29801v1 Announce Type: new Abstract: Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhi

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization

Model ReleasesDGX agent

arXiv:2605.29396v1 Announce Type: new Abstract: Safety alignment for large language models (LLMs) aims to reduce harmful or unsafe behavior while preserving general utility. However, recent findings r

Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization

Model ReleasesDGX agent

arXiv:2605.28969v1 Announce Type: cross Abstract: If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how f

← Previous
1…402403404405406…1053
Next →