AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,612 results
30 Apr 2026

Privacy-Preserving Federated Learning Framework for Distributed Chemical Process Optimization

Local AiDGX agent

arXiv:2604.26073v1 Announce Type: cross Abstract: Industrial chemical plants often operate under strict data confidentiality constraints, making centralized data-driven process modeling difficult. Fed

Reasoning Gets Harder for LLMs Inside A Dialogue

Model ReleasesDGX agent

arXiv:2603.20133v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that d

SciMDR: Advancing Scientific Multimodal Document Reasoning

Model ReleasesDGX agent

arXiv:2603.12249v2 Announce Type: replace-cross Abstract: Constructing scientific multimodal document reasoning datasets for foundation model training involves an inherent trade-off among scale, faith

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences

Model ReleasesDGX agent

arXiv:2509.11295v2 Announce Type: replace Abstract: Developing effective prompts demands significant cognitive investment to generate reliable, high-quality responses from Large Language Models (LLMs)

Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall

ResearchDGX agent

arXiv:2505.13963v3 Announce Type: replace Abstract: Quantization methods are widely used to accelerate inference and streamline the deployment of large language models (LLMs). Although quantization's

tweeted about this yesterday and Cursor already dropped the alpha today! 🚀very cool to see how us, them, and others have converged on good …

Model ReleasesDGX agent

tweeted about this yesterday and Cursor already dropped the alpha today! 🚀very cool to see how us, them, and others have converged on good design patterns in Agent + Harness Engineering: 1. Tuning dif

Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness

Model ReleasesDGX agent

arXiv:2512.03992v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are essential for embodied AI and safety-critical applications, such as robotics and autonomous systems. However

29 Apr 2026

Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

Model ReleasesDGX agent

arXiv:2604.25850v1 Announce Type: new Abstract: Harnesses have become a central determinant of coding-agent performance, shaping how models interact with repositories, tools, and execution environment

BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate

SafetyDGX agent

arXiv:2604.25203v1 Announce Type: new Abstract: Deploying guardrails for custom policies remains challenging, as generic safety models fail to capture task-specific requirements, while prompting LLMs

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding

Model ReleasesDGX agent

arXiv:2512.12087v3 Announce Type: replace Abstract: The growing demand for long-context inference capabilities in Large Language Models (LLMs) has intensified the computational and memory bottlenecks

Conditional Flow Matching for Probabilistic Downscaling of Maximum 3-day Snowfall in Alaska

ResearchDGX agent

arXiv:2604.25172v1 Announce Type: cross Abstract: Precipitation in complex terrain is governed by orographic processes operating at scales of a few kilometers, yet climate models typically run at reso

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

TutorialsDGX agent

arXiv:2604.25477v1 Announce Type: new Abstract: Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance t

DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference

ResearchDGX agent

arXiv:2510.19669v4 Announce Type: replace Abstract: Recent reasoning Large Language Models (LLMs) demonstrate remarkable problem-solving abilities but often generate long thinking traces whose utility

Doing More With Less: Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling

Model ReleasesDGX agent

arXiv:2604.25098v1 Announce Type: cross Abstract: While current Large Language Models (LLMs) exhibit remarkable reasoning capabilities through test-time compute scaling (TTS), their massive parameter

FED-FSTQ: Fisher-Guided Token Quantization for Communication-Efficient Federated Fine-Tuning of LLMs on Edge Devices

Model ReleasesDGX agent

arXiv:2604.25421v1 Announce Type: new Abstract: Federated fine-tuning provides a practical route to adapt large language models (LLMs) on edge devices without centralizing private data, yet in mobile

Golden RPG: Confidence-Adaptive Region-Aware Noise for Compositional Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2604.25314v1 Announce Type: new Abstract: Compositional text-to-image (T2I) generation requires a model to honour multiple sub-prompts that describe distinct image regions. Recent work shows tha

I released LLM 0.32a0 this morning, a major backwards-compatible refactor of my LLM Python library and CLI tool for working with language mo…

Model ReleasesDGX agent

I released LLM 0.32a0 this morning, a major backwards-compatible refactor of my LLM Python library and CLI tool for working with language models - the new changes should help LLM work better with reas

Improving LLM Predictions via Inter-Layer Structural Encoders

Model ReleasesDGX agent

arXiv:2603.22665v2 Announce Type: replace Abstract: The standard practice in Large Language Models (LLMs) is to base predictions on final-layer representations. However, intermediate layers encode com

Limited Linguistic Diversity in Embodied AI Datasets

ResearchDGX agent

arXiv:2601.03136v2 Announce Type: replace Abstract: Language plays a critical role in Vision-Language-Action (VLA) models, yet the linguistic characteristics of the datasets used to train and evaluate

LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation

Model ReleasesDGX agent

arXiv:2604.25665v1 Announce Type: new Abstract: Reliable evaluation of large language model (LLM)-generated summaries remains an open challenge, particularly across heterogeneous domains and document

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition

Model ReleasesDGX agent

arXiv:2512.07348v2 Announce Type: replace Abstract: In controllable image generation, synthesizing coherent and consistent images from multiple reference inputs, i.e., Multi-Image Composition (MICo),

Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study

Model ReleasesDGX agent

arXiv:2602.17262v2 Announce Type: replace Abstract: Human self-report questionnaires are increasingly used in NLP to benchmark and audit large language models (LLMs), from persona consistency to safet

RCProb: Probabilistic Rule Extraction for Efficient Simplification of Tree Ensembles

Model ReleasesDGX agent

arXiv:2604.25304v1 Announce Type: new Abstract: Tree ensembles are widely used in industrial machine learning due to their strong predictive performance and efficient training procedures. However, as

Relational In-Context Learning via Synthetic Pre-training with Structural Prior

ApplicationsDGX agent

arXiv:2603.03805v2 Announce Type: replace Abstract: Relational Databases (RDBs) are the backbone of modern business, yet they lack foundation models comparable to those in text or vision. A key obstac

ReSim: Reliable World Simulation for Autonomous Driving

SafetyDGX agent

arXiv:2506.09981v2 Announce Type: replace Abstract: How can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusivel

Sensitivity-Based Tube NMPC for Cooperative Aerial Structures Under Parametric Uncertainty

ResearchDGX agent

arXiv:2604.25766v1 Announce Type: new Abstract: This paper presents a sensitivity-based tube Nonlinear Model Predictive Control (NMPC) framework for cooperative aerial chains under bounded parametric

TopoMamba: Topology-Aware Scanning and Fusion for Segmenting Heterogeneous Medical Visual Media

ResearchDGX agent

arXiv:2604.25545v1 Announce Type: new Abstract: Visual state-space models (SSMs) have shown strong potential for medical image segmentation, yet their effectiveness is often limited by two practical i

When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs

Model ReleasesDGX agent

arXiv:2510.07499v2 Announce Type: replace Abstract: Recent Long-Context Language Models (LCLMs) can process hundreds of thousands of tokens in a single prompt, enabling new opportunities for knowledge

Xiaomi MiMo-V2.5-Pro achieves multiple breakthroughs in the latest Arena rankings (Apr 26, 2026) 🔥 🏆 Text Arena (Expert) — #6 globally | #…

ApplicationsDGX agent

Xiaomi MiMo-V2.5-Pro achieves multiple breakthroughs in the latest Arena rankings (Apr 26, 2026) 🔥 🏆 Text Arena (Expert) — #6 globally | #1 open-source model Also #1 among Chinese models, with Xiaomi

28 Apr 2026

Aligning with Your Own Voice: Self-Corrected Preference Learning for Hallucination Mitigation in LVLMs

SafetyDGX agent

arXiv:2604.24395v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) frequently suffer from hallucinations. Existing preference learning-based approaches largely rely on proprietary mo

amazing! it’s like talking to an ai from the past

AgentsDGX agent

amazing! it’s like talking to an ai from the past Announcing Talkie: a new, open-weight historical LLM! We trained and finetuned a 13B model on a newly-curated dataset of only pre-1930 data. Try it be

Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis

SafetyDGX agent

arXiv:2604.23072v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly tasked with complex real-world analysis (e.g., in financial forecasting, scientific discovery), yet t

AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation

Model ReleasesDGX agent

arXiv:2604.24086v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have been demonstrated possessing strong zero-shot generalization for robot control, their massive parameter

Benchmarking Source-Sensitive Reasoning in Turkish: Humans and LLMs under Evidential Trust Manipulation

Local AiDGX agent

arXiv:2604.24665v1 Announce Type: cross Abstract: This paper investigates whether source trustworthiness shapes Turkish evidential morphology and whether large language models (LLMs) track this sensit

Beyond Local vs. External: A Game-Theoretic Framework for Trustworthy Knowledge Acquisition

Model ReleasesDGX agent

arXiv:2604.23413v1 Announce Type: new Abstract: Cloud-hosted Large Language Models (LLMs) offer unmatched reasoning capabilities and dynamic knowledge, yet submitting raw queries to these external ser

BIR-Adapter: A parameter-efficient diffusion adapter for blind image restoration

Model ReleasesDGX agent

arXiv:2509.06904v3 Announce Type: replace Abstract: We introduce the BIR-Adapter, a parameter-efficient diffusion adapter for blind image restoration. Diffusion-based restoration methods have demonstr

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Model ReleasesDGX agent

arXiv:2604.23781v1 Announce Type: new Abstract: Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surroundi

Computational Design and Co-Robotic Fabrication for Material Reuse in Architecture

ApplicationsDGX agent

arXiv:2604.24648v1 Announce Type: new Abstract: Climate change and resource depletion demand a shift from the dominant linear 'take-make-use-dispose' paradigm of construction toward circular, low-wast

Continual Calibration: Coverage Can Collapse Before Accuracy in Lifelong LLM Fine-Tuning

ResearchDGX agent

arXiv:2604.23987v1 Announce Type: new Abstract: Continual learning for large language models is typically evaluated through accuracy retention under sequential fine-tuning. We argue that this perspect

Cortex-Inspired Continual Learning: Unsupervised Instantiation and Recovery of Functional Task Networks

Model ReleasesDGX agent

arXiv:2604.24637v1 Announce Type: cross Abstract: Block-sequential continual learning demands that a single model both protect prior solutions from catastrophic forgetting and efficiently infer at inf

Coverage-Based Calibration for Post-Training Quantization via Weighted Set Cover over Outlier Channels

Model ReleasesDGX agent

arXiv:2604.24008v1 Announce Type: new Abstract: Post-Training Quantization (PTQ) compresses large language models to low bit-widths using a small calibration set, and its quality depends strongly on w

DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting

Model ReleasesDGX agent

arXiv:2604.23968v1 Announce Type: cross Abstract: Accurate time series forecasting in scientific domains such as climate modeling, physiological monitoring, and energy systems benefits from both compe

Deep Learning of Solver-Aware Turbulence Closures from Nudged LES Dynamics

TutorialsDGX agent

arXiv:2604.23874v1 Announce Type: cross Abstract: Deep learning approaches have shown remarkable promise in turbulence closure modeling for large eddy simulations (LES). The differentiable physics par

Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing

Model ReleasesDGX agent

arXiv:2604.24162v1 Announce Type: cross Abstract: Defending against backdoor attacks in large language models remains a critical practical challenge. Existing defenses mitigate these threats but typic

Deploy DINO with Many-to-Many Association

Model ReleasesDGX agent

arXiv:2604.23670v1 Announce Type: new Abstract: Motivated by the limited generalization of supervised image matching models to unseen image domains, we explore the zero-shot deployment of DINO feature

DiffuSAM: Diffusion-Based Prompt-Free SAM2 for Few-Shot and Source-Free Medical Image Segmentation

ResearchDGX agent

arXiv:2604.24719v1 Announce Type: new Abstract: Segmentation models such as Segment Anything Model (SAM) and SAM2 achieve strong prompt-driven zero-shot performance. However, their training on natural

EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2604.23348v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and generation, and are increasingly used in

Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task

ApplicationsDGX agent

arXiv:2604.23730v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on legal benchmarks, including multiple-choice components of bar exams. However, their capaci

Explanation Quality Assessment as Ranking with Listwise Rewards

SafetyDGX agent

arXiv:2604.24176v1 Announce Type: new Abstract: We reformulate explanation quality assessment as a ranking problem rather than a generation problem. Instead of optimizing models to produce a single 'b

FlashOverlap: Minimizing Tail Latency in Communication Overlap for Distributed LLM Training

ResearchDGX agent

arXiv:2604.24013v1 Announce Type: cross Abstract: The rapid growth in the size of large language models has necessitated the partitioning of computational workloads across accelerators such as GPUs, T

Generating Place-Based Compromises Between Two Points of View

Model ReleasesDGX agent

arXiv:2604.24536v1 Announce Type: new Abstract: Large Language Models (LLMs) excel academically but struggle with social intelligence tasks, such as creating good compromises. In this paper, we presen

Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference

ResearchDGX agent

arXiv:2503.10666v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have become widely used across various domains spanning search engines, code generation, and text creation. Howev

How we built the most performant DeepSeek V3.2, MiniMax-M2.5 and Qwen 3.5 397B on DigitalOcean Serverless Inference

Model ReleasesDGX agent

DigitalOcean describes their optimization and deployment of three large language models (DeepSeek V3.2, MiniMax-M2.5, and Qwen 3.5 397B) on their Serverless Inference platform, likely leveraging NVIDI

Isotonic Layer: A Unified Framework for Recommendation Calibration and Debiasing

SafetyDGX agent

arXiv:2603.06589v2 Announce Type: replace-cross Abstract: Model calibration and debiasing are fundamental yet operationally expensive challenges in large-scale recommendation systems. Existing approac

KLong: Training LLM Agent for Extremely Long-horizon Tasks

Model ReleasesDGX agent

arXiv:2602.17547v3 Announce Type: replace Abstract: This paper introduces KLong, an open-source LLM agent trained to solve extremely long-horizon tasks. The principle is to first cold-start the model

MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG

Model ReleasesDGX agent

arXiv:2604.24564v1 Announce Type: new Abstract: Multimodal Retrieval-Augmented Generation (MRAG) addresses key limitations of Multimodal Large Language Models (MLLMs), such as hallucination and outdat

MEMCoder: Multi-dimensional Evolving Memory for Private-Library-Oriented Code Generation

Model ReleasesDGX agent

arXiv:2604.24222v1 Announce Type: cross Abstract: Large Language Models (LLMs) excel at general code generation, but their performance drops sharply in enterprise settings that rely on internal privat

MetaErr: Towards Predicting Error Patterns in Deep Neural Networks

Model ReleasesDGX agent

arXiv:2604.23289v1 Announce Type: cross Abstract: Due to the unprecedented success of deep learning, it has become an integral component in several multimedia computing applications in todays world. U

Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training

Model ReleasesDGX agent

arXiv:2506.20332v4 Announce Type: replace Abstract: Vision-language model-based mobile agents have gained the ability to understand complex instructions and mobile screenshots, benefiting from reinfor

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation

Model ReleasesDGX agent

arXiv:2604.23789v1 Announce Type: new Abstract: While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Fur

← Previous
1…359360361362363…1044
Next →