AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,573
  • Agents7,495
  • Applications5,364
  • Concepts5
  • Hardware1,813
  • Industry6,149
  • Local Ai4,892
  • Model Releases23,569
  • Research19,967
  • Safety13,263
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,573
  • Agents7,495
  • Applications5,364
  • Concepts5
  • Hardware1,813
  • Industry6,149
  • Local Ai4,892
  • Model Releases23,569
  • Research19,967
  • Safety13,263
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

87,573Total entries
1Added by human
87,572Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,952 results
14 Apr 2026

The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise

ResearchDGX agent

arXiv:2604.09780v1 Announce Type: new Abstract: Mixture of Experts (MoEs) are now ubiquitous in large language models, yet the mechanisms behind their 'expert specialization' remain poorly understood.

(This is a big part of what was called emergence in earlier academic work on unexpected LLM ability gains)

ApplicationsDGX agent

Ethan Mollick discusses the concept of 'emergence' in large language models (LLMs), referring to the phenomenon where AI systems appear to suddenly develop unexpected capabilities as they scale. The p

Toward Generalized Cross-Lingual Hateful Language Detection with Web-Scale Data and Ensemble LLM Annotations

Model Releases
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.09625v1 Announce Type: new Abstract: We study whether large-scale unlabelled web data and LLM-based synthetic annotations can improve multilingual hate speech detection. Starting from texts

Trajectory-based actuator identification via differentiable simulation

SafetyDGX agent

arXiv:2604.10351v1 Announce Type: new Abstract: Accurate actuation models are critical for bridging the gap between simulation and real robot behavior, yet obtaining high-fidelity actuator dynamics ty

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards

Model ReleasesDGX agent

arXiv:2604.10110v1 Announce Type: new Abstract: Large Language Models (LLMs) have become a key foundation for enabling personalized smart home experiences. While existing studies have explored how sma

Tuning Qwen2.5-VL to Improve Its Web Interaction Skills

Model ReleasesDGX agent

arXiv:2604.09571v1 Announce Type: cross Abstract: Recent advances in vision-language models (VLMs) have sparked growing interest in using them to automate web tasks, yet their feasibility as independe

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization

Model ReleasesDGX agent

arXiv:2604.10721v1 Announce Type: cross Abstract: Natural-language Guided Cross-view Geo-localization (NGCG) aims to retrieve geo-tagged satellite imagery using textual descriptions of ground scenes.

Variable Selection Using Relative Importance Rankings

Model ReleasesDGX agent

arXiv:2509.10853v2 Announce Type: replace-cross Abstract: Although conceptually related, variable selection and relative importance (RI) analysis have been treated quite differently in the literature.

Virtual Smart Metering in District Heating Networks via Heterogeneous Spatial-Temporal Graph Neural Networks

Model ReleasesDGX agent

arXiv:2604.10166v1 Announce Type: cross Abstract: Intelligent operation of thermal energy networks aims to improve energy efficiency, reliability, and operational flexibility through data-driven contr

VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites

ResearchDGX agent

arXiv:2506.14629v3 Announce Type: replace-cross Abstract: Mosquito-borne diseases pose a major global health risk, requiring early detection and proactive control of breeding sites to prevent outbreak

Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval

ResearchDGX agent

arXiv:2604.10167v1 Announce Type: cross Abstract: Multi-vector models dominate Visual Document Retrieval (VDR) due to their fine-grained matching capabilities, but their high storage and computational

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

SafetyDGX agent

arXiv:2510.26202v2 Announce Type: replace-cross Abstract: Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback d

13 Apr 2026

A Benchmark of Dexterity for Anthropomorphic Robotic Hands

Model ReleasesDGX agent

arXiv:2604.09294v1 Announce Type: new Abstract: Dexterity is a central yet ambiguously defined concept in the design and evaluation of anthropomorphic robotic hands. In practice, the term is often use

A Closer Look at the Application of Causal Inference in Graph Representation Learning

ResearchDGX agent

arXiv:2604.08890v1 Announce Type: cross Abstract: Modeling causal relationships in graph representation learning remains a fundamental challenge. Existing approaches often draw on theories and methods

Agentic Jackal: Live Execution and Semantic Value Grounding for Text-to-JQL

Model ReleasesDGX agent

arXiv:2604.09470v1 Announce Type: new Abstract: Translating natural language into Jira Query Language (JQL) requires resolving ambiguous field references, instance-specific categorical values, and com

Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography

Model ReleasesDGX agent

arXiv:2604.09066v1 Announce Type: new Abstract: Linguistic steganography based on language models typically assumes that steganographic texts are transmitted without alteration, making them fragile to

AnimaYume - Anima finetune.

Local AiDGX agent

AnimaYume is a text-to-image model fine-tuned from Anima, a 2-billion-parameter anime-focused image generation model developed by CircleStone Labs in collaboration with Comfy Org, which is itself buil

Biologically-Grounded Multi-Encoder Architectures as Developability Oracles for Antibody Design

Model ReleasesDGX agent

arXiv:2604.09369v1 Announce Type: cross Abstract: Generative models can now propose thousands of de novo antibody sequences, yet translating these designs into viable therapeutics remains constrained

CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Intervention

SafetyDGX agent

arXiv:2603.18561v2 Announce Type: replace Abstract: Planning-oriented end-to-end driving models show great promise, yet they fundamentally learn statistical correlations instead of true causal relatio

Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment

SafetyDGX agent

arXiv:2505.18600v3 Announce Type: replace-cross Abstract: Modern single-image super-resolution (SISR) models deliver photo-realistic results at the scale factors on which they are trained, but collaps

EinsteinArena is open-source and the leaderboard is live. We welcome contributions and feedback! → https://www.together.ai/blog/einsteinaren…

ToolsDGX agent

EinsteinArena is an open-source benchmarking platform developed by Together AI designed to evaluate and rank AI models, with a live public leaderboard tracking model performance. The project welcomes

Enterprises power agentic workflows in Cloudflare Agent Cloud with OpenAI

Model ReleasesDGX agent

Cloudflare and OpenAI have partnered to create the Cloudflare Agent Cloud, a platform that enables enterprises to build and deploy agentic AI workflows at scale using OpenAI's models and APIs. The int

Face vs body Zit Lora

Local AiDGX agent

This Reddit post from r/StableDiffusion discusses a community comparison or showcase involving a 'Zit LoRA' — a Stable Diffusion LoRA model — examining how it performs differently when applied to face

Gemini Robotics-ER 1.6: Powering real-world robotics tasks through enhanced embodied reasoning

Model ReleasesDGX agent

Gemini Robotics-ER 1.6, introduced by Google DeepMind, is a significant upgrade to their reasoning-first robotics model that specializes in visual and spatial understanding, task planning, and success

Generative View Stitching

TutorialsDGX agent

arXiv:2510.24718v3 Announce Type: replace Abstract: Autoregressive video diffusion models are capable of long rollouts that are stable and consistent with history, but they are unable to guide the cur

HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?

Model ReleasesDGX agent

arXiv:2604.09408v1 Announce Type: new Abstract: Frontier coding agents solve complex tasks when given complete context but collapse when specifications are incomplete or ambiguous. The bottleneck is n

Hitem3D 2.0: Multi-View Guided Native 3D Texture Generation

SafetyDGX agent

arXiv:2604.09231v1 Announce Type: new Abstract: Although recent advances have improved the quality of 3D texture generation, existing methods still struggle with incomplete texture coverage, cross-vie

Introducing DDTree: accelerates speculative decoding by drafting a tree with one block diffusion pass, then verifying multiple likely contin…

TutorialsDGX agent

Introducing DDTree: accelerates speculative decoding by drafting a tree with one block diffusion pass, then verifying multiple likely continuations together. Paper: https://liranringel.github.io/ddtre

Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation

Model ReleasesDGX agent

arXiv:2604.08797v1 Announce Type: cross Abstract: Stories are key to transmitting values across cultures, but their interpretation varies across linguistic and cultural contexts. Thus, we introduce mu

Mamba-Based Graph Convolutional Networks: Tackling Over-smoothing with Selective State Space

Model ReleasesDGX agent

arXiv:2501.15461v4 Announce Type: replace Abstract: Graph Neural Networks (GNNs) have shown great success in various graph-based learning tasks. However, it often faces the issue of over-smoothing as

Maybe hot take - I’ve read a bunch of RL for image generation papers over last few months and honestly it’s been pretty disappointing. All o…

TutorialsDGX agent

Maybe hot take - I’ve read a bunch of RL for image generation papers over last few months and honestly it’s been pretty disappointing. All of them are variations of GRPO and all of them are incrementa

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs

ResearchDGX agent

arXiv:2505.20638v2 Announce Type: replace-cross Abstract: While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music nec

NO MORE PAYING FOR API! NEW SOLUTION!

Local AiDGX agent

A Reddit post from the r/ollama community discussing a free alternative to paid AI API services, likely centered around using Ollama to run large language models locally. The post probably highlights

Ollama / Mistral with MCP to Mempalace

Model ReleasesDGX agent

This Reddit post from r/ollama discusses integrating Ollama-served Mistral with MemPalace — a free, locally-run AI memory system — via the Model Context Protocol (MCP). MemPalace runs entirely on a us

PhysInOne: Visual Physics Learning and Reasoning in One Suite

Model ReleasesDGX agent

arXiv:2604.09415v1 Announce Type: cross Abstract: We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike exi

PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos

Model ReleasesDGX agent

arXiv:2604.08991v1 Announce Type: cross Abstract: Small object-centric spatial understanding in indoor videos remains a significant challenge for multimodal large language models (MLLMs), despite its

Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance

Model ReleasesDGX agent

arXiv:2604.08881v1 Announce Type: new Abstract: In real-world deployments, Vision-Language Large Models (VLLMs) face critical challenges from multilingual and multimodal composite attacks: harmful ima

Relational Visual Similarity

ApplicationsDGX agent

arXiv:2512.07833v2 Announce Type: replace-cross Abstract: Humans do not just see attribute similarity -- we also see relational similarity. An apple is like a peach because both are reddish fruit, but

RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation

Model ReleasesDGX agent

arXiv:2510.17640v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable performance on complex tasks through imitation learning in recent robotic man

SenBen: Sensitive Scene Graphs for Explainable Content Moderation

Model ReleasesDGX agent

arXiv:2604.08819v1 Announce Type: cross Abstract: Content moderation systems classify images as safe or unsafe but lack spatial grounding and interpretability: they cannot explain what sensitive behav

SHIFT: Steering Hidden Intermediates in Flow Transformers

SafetyDGX agent

arXiv:2604.09213v1 Announce Type: new Abstract: Diffusion models have become leading approaches for high-fidelity image generation. Recent DiT-based diffusion models, in particular, achieve strong pro

SSPO: Subsentence-level Policy Optimization

SafetyDGX agent

arXiv:2511.04256v2 Announce Type: replace Abstract: As a key component of large language model (LLM) post-training, Reinforcement Learning from Verifiable Rewards (RLVR) has substantially improved rea

Think Less, Know More: State-Aware Reasoning Compression with Knowledge Guidance for Efficient Reasoning

SafetyDGX agent

arXiv:2604.09150v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks by leveraging long Chain-of-Thought (CoT), but often suffer from overthinking,

TinyNeRV: Compact Neural Video Representations via Capacity Scaling, Distillation, and Low-Precision Inference

Model ReleasesDGX agent

arXiv:2604.09220v1 Announce Type: new Abstract: Implicit neural video representations encode entire video sequences within the parameters of a neural network and enable constant time frame reconstruct

Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments

Model ReleasesDGX agent

arXiv:2604.09038v1 Announce Type: cross Abstract: Robust geo-localization in changing environmental conditions is critical for long-term aerial autonomy. While visual place recognition (VPR) models pe

Unified Multimodal Uncertain Inference

Model ReleasesDGX agent

arXiv:2604.08701v1 Announce Type: new Abstract: We introduce Unified Multimodal Uncertain Inference (UMUI), a multimodal inference task spanning text, audio, and video, where models must produce calib

V-CAGE: Vision-Closed-Loop Agentic Generation Engine for Robotic Manipulation

AgentsDGX agent

arXiv:2604.09036v1 Announce Type: new Abstract: Scaling Vision-Language-Action (VLA) models requires massive datasets that are both semantically coherent and physically feasible. However, existing sce

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

SafetyDGX agent

arXiv:2604.09330v1 Announce Type: cross Abstract: Recent advances in robot foundation models trained on large-scale human teleoperation data have enabled robots to perform increasingly complex real-wo

VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images

ResearchDGX agent

arXiv:2604.09531v1 Announce Type: cross Abstract: Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition. One plausible contr

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures

Model ReleasesDGX agent

arXiv:2604.09048v1 Announce Type: cross Abstract: While the large energy consumption of Large Language Models (LLMs) is recognized by the community, system operators lack guidance for energy-efficient

we used Chandra-OCR-2 by @datalabto: https://huggingface.co/datalab-to/chandra-ocr-2 Full write-up by @NielsRogge: https://huggingface.co/bl…

IndustryDGX agent

Chandra-OCR-2 is an optical character recognition model developed by DataLab, available on Hugging Face at datalab-to/chandra-ocr-2. The model was highlighted by Hugging Face CEO Clément Delangue, wit

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning

Model ReleasesDGX agent

arXiv:2510.07517v5 Announce Type: replace Abstract: Multi-agent debate (MAD) aims to improve large language model (LLM) reasoning by letting multiple agents exchange answers and then aggregate their o

12 Apr 2026

Can't get a good coding setup on Macbook Pro M3 Max 36GB

Local AiDGX agent

This Reddit thread from r/ollama discusses a user's difficulty achieving a satisfactory local AI coding assistant setup using Ollama on a MacBook Pro M3 Max with 36GB of unified memory. The discussion

Gemma 4 audio with MLX

Model ReleasesDGX agent

Thanks to a tip from Rahim Nathwani, here's a uv run recipe for transcribing an audio file on macOS using the 10.28 GB Gemma 4 E2B model with MLX and mlx-vlm: uv run --python 3.13 --with mlx_vlm --wit

Greg Rutkowski Anima Lora from Circlestone Labs (Anima makers) with training params

Local AiDGX agent

This Reddit post discusses a LoRA trained to emulate the style of digital artist Greg Rutkowski, built on top of Circlestone Labs' Anima model — a 2 billion parameter text-to-image model created via a

Most people building AI agents don't realize this until it's too late. Your agent's memory isn't a feature. It's your moat. 1. Closed harnes…

Model ReleasesDGX agent

Most people building AI agents don't realize this until it's too late. Your agent's memory isn't a feature. It's your moat. 1. Closed harness = they own your memory 2. Switch models → lose all context

Tile upscale controlnet with Z-Image-Base? Has anybody achieved good results?

Local AiDGX agent

This Reddit thread from r/StableDiffusion discusses community experiences using ControlNet Tile upscaling in combination with Z-Image-Base, a Stable Diffusion base model. ControlNet Tile models are of

Topping the charts!

AgentsDGX agent

Nous Research shared a post celebrating a model or benchmark achievement reaching the top of a performance leaderboard, likely referencing one of their Hermes or other open-source model releases. The

Z-Image Turbo Checkpoint - Deedeemegadoodo Edition

Local AiDGX agent

The 'Z-Image Turbo Checkpoint - Deedeemegadoodo Edition' is a community-shared checkpoint on r/StableDiffusion based on Z-Image Turbo, a distilled version of Z-Image, a 6B image model developed by the

11 Apr 2026

fine-tune LTX 2.3 with his own dataset?

Local AiDGX agent

This r/StableDiffusion thread discusses how to fine-tune the LTX-Video 2.3 model on a personal dataset, with the primary approach being LoRA (Low-Rank Adaptation), which fine-tunes a large AI model on

← Previous
1…370371372373374…1050
Next →