AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,565 results
15 May 2026

A Hormone-inspired Emotion Layer for Transformer language models (HELT)

ResearchDGX agent

arXiv:2605.13858v1 Announce Type: cross Abstract: Large Language Models have demonstrated remarkable capabilities in generating contextually relevant and grammatically correct text. However, they fund

Causal Foundation Models with Continuous Treatments

ResearchDGX agent

arXiv:2605.15133v1 Announce Type: new Abstract: Causal inference, estimating causal effects from observational data, is a fundamental tool in many disciplines. Of particular importance across a variet

Conditional Attribute Estimation with Autoregressive Sequence Models

Local AiDGX agent

arXiv:2605.14004v1 Announce Type: new Abstract: Generative models are often trained with a next-token prediction objective, yet many downstream applications require the ability to estimate or control

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

EMA: Efficient Model Adaptation for Learning-based Systems

HardwareDGX agent

arXiv:2605.13942v1 Announce Type: new Abstract: Machine learning (ML) is increasingly applied to optimize system performance in tasks such as resource management and network simulation. Unlike traditi

Generalizing Score-based generative models for Heavy-tailed Distributions

ResearchDGX agent

arXiv:2603.00772v2 Announce Type: replace-cross Abstract: Score-based generative models (SGMs) have achieved remarkable empirical success, motivating their application to a broad range of data distrib

HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling

SafetyDGX agent

arXiv:2605.14877v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models have recently demonstrated impressive image generation quality while maintaining low latency. However, they suffer fr

MeMo: Memory as a Model

ApplicationsDGX agent

arXiv:2605.15156v1 Announce Type: cross Abstract: Large language models (LLMs) achieve strong performance across a wide range of tasks, but remain frozen after pretraining until subsequent updates. Ma

Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey

Model ReleasesDGX agent

arXiv:2605.13919v1 Announce Type: new Abstract: Multilingual knowledge editing (MKE) remains challenging because language-specific edits interfere with one another, even when locate-then-edit methods

Multi-Dimensional Model Integrity and Responsibility Assessment Index and Scoring Framework

SafetyDGX agent

arXiv:2605.14550v1 Announce Type: new Abstract: Artificial intelligence in high-stakes tabular domains cannot be evaluated by predictive performance alone, yet current practice still assesses explaina

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

Local AiDGX agent

arXiv:2603.02115v2 Announce Type: replace-cross Abstract: General-purpose robot reward models are typically trained to predict absolute task progress from expert demonstrations, providing only local,

TopoPrimer: The Missing Topological Context in Forecasting Models

ResearchDGX agent

arXiv:2605.15035v1 Announce Type: new Abstract: We introduce TopoPrimer, a framework that makes the global topological structure of the series population an explicit input to any forecasting model. To

Towards Fine-Grained and Verifiable Concept Bottleneck Models

Local AiDGX agent

arXiv:2605.14210v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) offer interpretable alternatives to black-box predictors by introducing human-relatable concepts before the final out

14 May 2026

BEHAVE: A Hybrid AI Framework for Real-Time Modeling of Collective Human Dynamics

SafetyDGX agent

arXiv:2605.12730v1 Announce Type: new Abstract: Existing AI systems for modeling human behavior operate at the level of individuals or detect events after they occur. As a result, they systematically

Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models

ResearchDGX agent

arXiv:2605.12519v1 Announce Type: cross Abstract: Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifiable rewards

Decoupled and Divergence-Conditioned Prompt for Multi-domain Dynamic Graph Foundation Models

ApplicationsDGX agent

arXiv:2605.13540v1 Announce Type: cross Abstract: Dynamic graphs are ubiquitous in real-world systems, and building generalizable dynamic Graph Foundation Models has become a frontier in graph learnin

deepagents v0.6 is our biggest release yet!!! it’s all about perf - at the model layer w harness profiles, agent layer w code interpreter, a…

AgentsDGX agent

deepagents v0.6 is our biggest release yet!!! it’s all about perf - at the model layer w harness profiles, agent layer w code interpreter, and at scale w streaming and delta channels context hub backe

Efficient Generative Prediction for EHR Foundation Models: The SCOPE and REACH Estimators

ApplicationsDGX agent

arXiv:2602.03730v2 Announce Type: replace-cross Abstract: Generative foundation models trained on tokenized electronic health record (EHR) timelines show promise for clinical outcome prediction via Mo

KamonBench: A Grammar-Based Dataset for Evaluating Compositional Factor Recovery in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.13322v1 Announce Type: new Abstract: Kamon (family crests) are an important part of Japanese culture and a natural test case for compositional visual recognition: each crest combines a smal

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

SafetyDGX agent

arXiv:2605.12729v1 Announce Type: cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), includ

Modeling Heterophily in Multiplex Graphs: An Adaptive Approach for Node Classification

TutorialsDGX agent

arXiv:2605.12699v1 Announce Type: cross Abstract: Existing multiplex graph models often assume homophily, where connected nodes tend to belong to the same class or share similar attributes. Consequent

Plan for Speed: Dilated Scheduling for Masked Diffusion Language Models

ResearchDGX agent

arXiv:2506.19037v4 Announce Type: replace-cross Abstract: Masked diffusion language models (MDLMs) promise fast, non-autoregressive text generation, yet existing samplers, which pick tokens to unmask

Predict-Project-Renoise: Sampling Diffusion Models under Hard Constraints

ResearchDGX agent

arXiv:2601.21033v2 Announce Type: replace Abstract: Diffusion models cannot enforce hard constraints, yet applications in the physical sciences demand exact satisfaction of conservation laws, boundary

Proximal-Based Generative Modeling for Bayesian Inverse Problems

SafetyDGX agent

arXiv:2605.13278v1 Announce Type: cross Abstract: Score-based diffusion models demonstrate superior performance in generative tasks but encounter fundamental bottlenecks in inverse problems due to the

Together AI STT models now hold the top two spots for transcription speed on the @ArtificialAnlys Speech to Text leaderboard. NVIDIA Parakee…

HardwareDGX agent

Together AI STT models now hold the top two spots for transcription speed on the @ArtificialAnlys Speech to Text leaderboard. NVIDIA Parakeet TDT 0.6B V3 on Together AI ranks #1, transcribing 303 seco

Topo-R1: Detecting Topological Anomalies via Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.13054v2 Announce Type: replace Abstract: Topology is critical in tubular structures such as blood vessels, nerve fibers, and road networks, where connectivity and loop structure govern down

Training Large Language Models to Predict Clinical Events

Model ReleasesDGX agent

arXiv:2605.12817v1 Announce Type: cross Abstract: Longitudinal clinical notes contain rich evidence of how patients evolve over time, but converting this signal into training supervision for clinical

Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization

ResearchDGX agent

arXiv:2605.12756v1 Announce Type: cross Abstract: Large language models (LLMs) are pretrained by minimizing the cross-entropy loss for next-token prediction. In this paper, we study whether this optim

What’s the best model to use with RAG to create a locally hosted survival and off grid LLm?

Local AiDGX agent

This discussion explores which language models work best when combined with RAG (Retrieval-Augmented Generation) for building a locally hosted LLM focused on survival and off-grid living topics. The t

13 May 2026

A Mixture Autoregressive Image Generative Model on Quadtree Regions for Gaussian Noise Removal via Variational Bayes and Gradient Methods

ResearchDGX agent

arXiv:2605.11585v1 Announce Type: new Abstract: This paper addresses the problem of image denoising for grayscale images. We propose a probabilistic image generative model that combines a quadtree reg

BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models

HardwareDGX agent

arXiv:2512.12131v2 Announce Type: replace Abstract: The scale of transformer model pre-training is constrained by the increasing computation and communication cost. Low-rank bottleneck architectures o

Can Nano Banana 2 Replace Traditional Image Restoration Models? An Evaluation of Its Performance on Image Restoration Tasks

ResearchDGX agent

arXiv:2604.03061v2 Announce Type: replace Abstract: Recent advances in generative AI raise the question of whether general-purpose image editing models can serve as unified solutions for image restora

CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration

Model ReleasesDGX agent

arXiv:2605.11186v1 Announce Type: new Abstract: Auto-regressive decoding in Large Language Models (LLMs) is inherently memory-bound: every generation step requires loading the model weights and interm

Combining On-Policy Optimization and Distillation for Long-Context Reasoning in Large Language Models

SafetyDGX agent

arXiv:2605.12227v1 Announce Type: new Abstract: Adapting large language models (LLMs) to long-context tasks requires post-training methods that remain accurate and coherent over thousands of tokens. E

Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification

ResearchDGX agent

arXiv:2509.20899v3 Announce Type: replace Abstract: Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but exte

Correcting Selection Bias in Sparse User Feedback for Large Language Model Quality Estimation: A Multi-Agent Hierarchical Bayesian Approach

Model ReleasesDGX agent

arXiv:2605.12177v1 Announce Type: new Abstract: [Abridged] Production LLM deployments receive feedback from a non-random fraction of users: thumbs sit mostly in the tails of the satisfaction distribut

Dynamic Execution Commitment of Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2605.11567v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models predominantly adopt action chunking, i.e., predicting and committing to a short horizon of consecutive low-level act

From Model Uncertainty to Human Attention: Localization-Aware Visual Cues for Scalable Annotation Review

ResearchDGX agent

arXiv:2605.12303v1 Announce Type: cross Abstract: High-quality labeled data is essential for training robust machine learning models, yet obtaining annotations at scale remains expensive. AI-assisted

Introducing AutoScientist. Most model training fails outside of frontier labs. AutoScientist automates the full research loop so it doesn't …

ToolsDGX agent

AutoScientist is an AI system developed by Together AI that automates the complete research workflow to enable model training and scientific discovery outside of well-resourced frontier laboratories.

Is Monotonic Sampling Necessary in Diffusion Models?

ResearchDGX agent

arXiv:2605.11773v1 Announce Type: new Abstract: Diffusion models generate samples by iteratively denoising a Gaussian prior, traversing a sequence of noise levels that, in every published sampler, dec

One-Step Generative Modeling via Wasserstein Gradient Flows

ResearchDGX agent

arXiv:2605.11755v1 Announce Type: cross Abstract: Diffusion models and flow-based methods have shown impressive generative capability, especially for images, but their sampling is expensive because it

ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models

SafetyDGX agent

arXiv:2605.12446v1 Announce Type: cross Abstract: Large language models (LLMs) often produce answers with high certainty even when they are incorrect, making reliable confidence estimation essential f

Overparametrized models with posterior drift

ResearchDGX agent

arXiv:2506.23619v2 Announce Type: replace-cross Abstract: This paper investigates the impact of posterior drift on out-of-sample forecasting accuracy in overparametrized machine learning models. We do

Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling

AgentsDGX agent

arXiv:2605.12411v1 Announce Type: cross Abstract: AI agents negotiate and transact in natural language with unfamiliar counterparts: a buyer bot facing an unknown seller, or a procurement assistant ne

Pretraining Exposure Explains Popularity Judgments in Large Language Models

SafetyDGX agent

arXiv:2605.12382v1 Announce Type: new Abstract: Large language models (LLMs) exhibit systematic preferences for well-known entities, a phenomenon often attributed to popularity bias. However, the exte

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation

SafetyDGX agent

arXiv:2603.15759v2 Announce Type: replace-cross Abstract: Robot learning requires adaptation methods that improve reliably from limited, mixed-quality interaction data. This is especially challenging

Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models

ResearchDGX agent

arXiv:2605.10971v1 Announce Type: cross Abstract: Discrete diffusion language models (DLMs) generate text by iteratively denoising all positions in parallel, offering an alternative to autoregressive

What makes a word hard to learn? Modeling L1 influence on English vocabulary difficulty

TutorialsDGX agent

arXiv:2605.12281v1 Announce Type: new Abstract: What makes a word difficult to learn, and how does the difficulty depend on the learner's native language? We computationally model vocabulary difficult

12 May 2026

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models

SafetyDGX agent

arXiv:2605.08513v1 Announce Type: cross Abstract: Safety alignment in language models operates through two mechanistically distinct systems: refusal neurons that gate whether harmful knowledge is expr

Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study

ResearchDGX agent

arXiv:2605.09622v1 Announce Type: cross Abstract: Voxel-wise dose prediction is a critical yet challenging task in practical radiotherapy (RT) planning, as bespoke models trained from scratch often st

BaLoRA: Bayesian Low-Rank Adaptation of Large Scale Models

ResearchDGX agent

arXiv:2605.08110v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become the standard for fine-tuning large pre-trained models at reduced computational cost. However, its low-rank point

BetaEdit: Null-Space Constrained Sequential Model Editing

ResearchDGX agent

arXiv:2605.09285v1 Announce Type: new Abstract: Null-space-based methods have garnered considerable attention in model editing by constraining updates to the null space of the pre-existing knowledge r

Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction

Local AiDGX agent

arXiv:2605.08276v1 Announce Type: new Abstract: Cell-level dense prediction is central to computational pathology, but remains challenging due to fine-grained histological structures, strong domain sh

BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization

Model ReleasesDGX agent

arXiv:2605.10345v1 Announce Type: new Abstract: Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization

Compute Where it Counts: Self Optimizing Language Models

SafetyDGX agent

arXiv:2605.10875v1 Announce Type: cross Abstract: Efficient LLM inference research has largely focused on reducing the cost of each decoding step (e.g., using quantization, pruning, or sparse attentio

Continuum Robot Modeling with Action Conditioned Flow Matching

TutorialsDGX agent

arXiv:2605.09216v1 Announce Type: new Abstract: Predicting the shape of tendon driven continuum robots (TDCRs) at steady state from actuation remains challenging due to continuous deformation, complex

CORTEG: Foundation Models Enable Cross-Modality Representation Transfer from Scalp to Intracranial Brain Recordings

Model ReleasesDGX agent

arXiv:2605.10337v1 Announce Type: new Abstract: Intracranial electrocorticography (ECoG) offers high-signal-to-noise access to cortical activity for brain-computer interfaces, yet limited per-patient

Data-driven Circuit Discovery for Interpretability of Language Models

Local AiDGX agent

arXiv:2605.09129v1 Announce Type: new Abstract: Circuit discovery aims to explain how language models (LMs) implement a specific task by localizing and interpreting a circuit, a computational subgraph

DetRefiner: Model-Agnostic Detection Refinement with Feature Fusion Transformer

Local AiDGX agent

arXiv:2605.10190v1 Announce Type: new Abstract: Open-vocabulary object detection (OVOD) aims to detect both seen and unseen categories, yet existing methods often struggle to generalize to novel objec

Do Foundation Model Embeddings Improve Cross-Country Crop Yield Generalisation? A Leave-One-Country-Out Evaluation in Sub-Saharan Africa

Model ReleasesDGX agent

arXiv:2605.08113v1 Announce Type: cross Abstract: Accurate predictions of smallholder maize yields across national boundaries are critical for food security planning in sub-Saharan Africa, yet most pu

EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation

AgentsDGX agent

arXiv:2603.09465v3 Announce Type: replace-cross Abstract: Vision-Language-Action models have shown great promise for autonomous driving, yet they suffer from degraded perception after unfreezing the v

← Previous
1…160161162163164…1010
Next →