AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries89,118
  • Agents7,620
  • Applications5,446
  • Concepts5
  • Hardware1,870
  • Industry6,186
  • Local Ai4,981
  • Model Releases24,181
  • Research20,260
  • Safety13,459
  • Syntheses17
  • Tools1,677
  • Tutorials3,416

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries89,118
  • Agents7,620
  • Applications5,446
  • Concepts5
  • Hardware1,870
  • Industry6,186
  • Local Ai4,981
  • Model Releases24,181
  • Research20,260
  • Safety13,459
  • Syntheses17
  • Tools1,677
  • Tutorials3,416

Source
HumanDGX agent

89,118Total entries
1Added by human
89,117Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
64,222 results
14 Apr 2026

Engineering Resource-constrained Software Systems with DNN Components: a Concept-based Pruning Approach

TutorialsDGX agent

arXiv:2604.09988v1 Announce Type: cross Abstract: Deep Neural Networks (DNNs) are widely used by engineers to solve difficult problems that require predictive modeling from data. However, these models

EviRCOD: Evidence-Guided Probabilistic Decoding for Referring Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2604.10894v1 Announce Type: new Abstract: Referring Camouflaged Object Detection (Ref-COD) focuses on segmenting specific camouflaged targets in a query image using category-aligned references.

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE

ResearchDGX agent

arXiv:2603.06003v2 Announce Type: replace Abstract: Sparse Mixture-of-Experts (SMoE) language models achieve strong capability at low per-token compute, yet deployment remains constrained by memory fo

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Exploring Cross-Modal Flows for Few-Shot Learning

Model ReleasesDGX agent

arXiv:2510.14543v4 Announce Type: replace Abstract: Aligning features from different modalities, is one of the most fundamental challenges for cross-modal tasks. Although pre-trained vision-language m

Exploring the best way for UAV visual localization under Low-altitude Multi-view Observation Condition: a Benchmark

Model ReleasesDGX agent

arXiv:2503.10692v2 Announce Type: replace Abstract: Absolute Visual Localization (AVL) enables an Unmanned Aerial Vehicle (UAV) to determine its position in GNSS-denied environments by establishing ge

Face Density as a Proxy for Data Complexity: Quantifying the Hardness of Instance Count

SafetyDGX agent

arXiv:2604.09689v1 Announce Type: cross Abstract: Machine learning progress has historically prioritized model-centric innovations, yet achievable performance is frequently capped by the intrinsic com

FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning

SafetyDGX agent

arXiv:2604.10693v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has improved LLM reasoning, but models often generate explanations that appear coherent while containing unfaithful int

FAITH: Factuality Alignment through Integrating Trustworthiness and Honestness

SafetyDGX agent

arXiv:2604.10189v1 Announce Type: new Abstract: Large Language Models (LLMs) can generate factually inaccurate content even if they have corresponding knowledge, which critically undermines their reli

Forget about VAEs? SenseNova's NEO-unify achieves 31.5 PSNR without an encoder – Native Image Gen is coming.

Local AiDGX agent

SenseNova's NEO-unify is an encoder-free unified multimodal model built on a Mixture-of-Transformers (MoT) backbone that eliminates the need for traditional VAE encoders in image generation. The 2B NE

From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning

ResearchDGX agent

arXiv:2604.10517v1 Announce Type: new Abstract: Modern vision-language models achieve strong performance in static perception, but remain limited in the complex spatiotemporal reasoning required for e

GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO

ResearchDGX agent

arXiv:2601.06767v2 Announce Type: replace-cross Abstract: We present a Bengali mathematical reasoning model called GanitLLM (named after the Bangla word for mathematics, 'Ganit'), together with a new

GazeVaLM: A Multi-Observer Eye-Tracking Benchmark for Evaluating Clinical Realism in AI-Generated X-Rays

Model ReleasesDGX agent

arXiv:2604.11653v1 Announce Type: new Abstract: We introduce GazeVaLM, a public eye-tracking dataset for studying clinical perception during chest radiograph authenticity assessment. The dataset compr

GenProve: Learning to Generate Text with Fine-Grained Provenance

SafetyDGX agent

arXiv:2601.04932v2 Announce Type: replace Abstract: Large language models (LLM) often hallucinate, and while adding citations is a common solution, it is frequently insufficient for accountability as

Governed Reasoning for Institutional AI

Model ReleasesDGX agent

arXiv:2604.10658v1 Announce Type: new Abstract: Institutional decisions -- regulatory compliance, clinical triage, prior authorization appeal -- require a different AI architecture than general-purpos

GPT vs Claude in a bomberman-style 1v1 game

Model ReleasesDGX agent

A Reddit post on r/ChatGPT in which a user built or showcased a Bomberman-style 1v1 game pitting GPT (OpenAI) against Claude (Anthropic) as autonomous AI players, likely using their respective APIs to

HeceTokenizer: A Syllable-Based Tokenization Approach for Turkish Retrieval

Model ReleasesDGX agent

arXiv:2604.10665v1 Announce Type: new Abstract: HeceTokenizer is a syllable-based tokenizer for Turkish that exploits the deterministic six-pattern phonological structure of the language to construct

How LLMs Might Think

ResearchDGX agent

arXiv:2604.09674v1 Announce Type: new Abstract: Do large language models (LLMs) think? Daniel Stoljar and Zhihe Vincent Zhang have recently developed an argument from rationality for the claim that LL

How You Ask Matters! Adaptive RAG Robustness to Query Variations

Model ReleasesDGX agent

arXiv:2604.10745v1 Announce Type: new Abstract: Adaptive Retrieval-Augmented Generation (RAG) promises accuracy and efficiency by dynamically triggering retrieval only when needed and is widely used i

https://x.com/nousresearch/status/2043969403247616478?s=46

ResearchDGX agent

Nous Research shared an announcement or update on their X (formerly Twitter) account, likely relating to their ongoing work in AI model development, fine-tuning, or research releases. Nous Research is

I built a tool that uses diffusion to create user interfaces.

Local AiDGX agent

A community developer shared a custom-built tool on r/StableDiffusion that leverages diffusion model technology — typically used for AI image generation — to generate user interface (UI) designs or co

I Can't Believe TTA Is Not Better: When Test-Time Augmentation Hurts Medical Image Classification

Model ReleasesDGX agent

arXiv:2604.09697v1 Announce Type: cross Abstract: Test-time augmentation (TTA)--aggregating predictions over multiple augmented copies of a test input--is widely assumed to improve classification accu

I Walk the Line: Examining the Role of Gestalt Continuity in Object Binding for Vision Transformers

ResearchDGX agent

arXiv:2604.09942v1 Announce Type: cross Abstract: Object binding is a foundational process in visual cognition, during which low-level perceptual features are joined into object representations. Bindi

IMPACT: A Dataset for Multi-Granularity Human Procedural Action Understanding in Industrial Assembly

Model ReleasesDGX agent

arXiv:2604.10409v1 Announce Type: cross Abstract: We introduce IMPACT, a synchronized five-view RGB-D dataset for deployment-oriented industrial procedural understanding, built around real assembly an

Improving understanding and trust in AI: How users benefit from interval-based counterfactual explanations

ResearchDGX agent

arXiv:2604.09573v1 Announce Type: cross Abstract: Experimental user studies evaluating the effectiveness of different subtypes of post-hoc explanations for black-box models are largely nonexistent. Th

Inertial Magnetic SLAM Systems Using Low-Cost Sensors

Local AiDGX agent

arXiv:2512.10128v2 Announce Type: replace Abstract: Spatially inhomogeneous magnetic fields offer a valuable, non-visual information source for positioning. Among systems leveraging this, magnetic fie

Investigating Bias and Fairness in Appearance-based Gaze Estimation

Model ReleasesDGX agent

arXiv:2604.10707v1 Announce Type: new Abstract: While appearance-based gaze estimation has achieved significant improvements in accuracy and domain adaptation, the fairness of these systems across dif

KL Divergence Between Gaussians: A Step-by-Step Derivation for the Variational Autoencoder Objective

TutorialsDGX agent

arXiv:2604.11744v1 Announce Type: new Abstract: Kullback-Leibler (KL) divergence is a fundamental concept in information theory that quantifies the discrepancy between two probability distributions. I

KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation

ResearchDGX agent

arXiv:2509.20128v2 Announce Type: replace-cross Abstract: Audio-driven facial animation has made significant progress in multimedia applications, with diffusion models showing strong potential for tal

Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation

ApplicationsDGX agent

arXiv:2604.09955v1 Announce Type: new Abstract: Video Unsupervised Domain Adaptation (VUDA) poses a significant challenge in action recognition, requiring the adaptation of a model from a labeled sour

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs

AgentsDGX agent

arXiv:2603.27494v2 Announce Type: replace-cross Abstract: To enhance the perception and reasoning capabilities of multimodal large language models in complex visual scenes, recent research has introdu

LLM-as-Judge on a Budget

SafetyDGX agent

arXiv:2602.15481v2 Announce Type: replace Abstract: LLM-as-a-judge has emerged as a cornerstone technique for evaluating large language models by leveraging LLM reasoning to score prompt-response pair

LLM Nepotism in Organizational Governance

SafetyDGX agent

arXiv:2604.09620v1 Announce Type: cross Abstract: Large language models are increasingly used to support organizational decisions from hiring to governance, raising fairness concerns in AI-assisted ev

LTX 2.3 Lora Training - Data Set Captioning

Local AiDGX agent

This Reddit thread from r/StableDiffusion discusses best practices for captioning training datasets when fine-tuning LoRA models on LTX 2.3, Lightricks' video generation model. The community emphasis

MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization

SafetyDGX agent

arXiv:2601.07208v2 Announce Type: replace-cross Abstract: Group-Relative Policy Optimization (GRPO) has emerged as an efficient paradigm for aligning Large Language Models (LLMs), yet its efficacy is

MARLIN: Multi-Agent Reinforcement Learning Guided by Language-Based Inter-Robot Negotiation

SafetyDGX agent

arXiv:2410.14383v4 Announce Type: replace Abstract: Multi-agent reinforcement learning is a key method for training multi-robot systems. Through rewarding or punishing robots over a series of episodes

MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis

Model ReleasesDGX agent

arXiv:2604.11188v1 Announce Type: cross Abstract: Synthesizing high-quality mathematical reasoning data without human priors remains a significant challenge. Current approaches typically rely on seed

MAVEN-T: Multi-Agent enVironment-aware Enhanced Neural Trajectory predictor with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.10169v1 Announce Type: new Abstract: Trajectory prediction remains a critical yet challenging component in autonomous driving systems, requiring sophisticated reasoning capabilities while m

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval

Model ReleasesDGX agent

arXiv:2604.09552v1 Announce Type: cross Abstract: Engineering rulebooks and technical standards contain multimodal information like dense text, tables, and illustrations that are challenging for retri

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

SafetyDGX agent

arXiv:2604.09757v1 Announce Type: cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely te

Most AI assistants wait for you to ask. But a truly useful agent should notice you need help before you say anything. New research takes a s…

Model ReleasesDGX agent

Most AI assistants wait for you to ask. But a truly useful agent should notice you need help before you say anything. New research takes a serious shot at building proactive agents that work in real t

Navigating the generative AI journey: The Path-to-Value framework from AWS

ApplicationsDGX agent

The AWS Generative AI Path-to-Value (P2V) framework is a structured mental model and practical guide designed to help organizations move generative AI initiatives from ideation and experimentation thr

Openclaw with Gemma4 26B extremely slow and forget stuff

Local AiDGX agent

Users in the r/ollama community report that running OpenClaw with the Gemma4 26B model via Ollama results in pathologically slow first-turn performance, with the degree of slowdown scaling with OpenCl

Our technical report, including how we designed the reward and the performance/latency Pareto frontier, is live here: https://cognition.ai/b…

AgentsDGX agent

Cognition AI has published a technical report detailing the design methodology behind their reward system, likely related to their Devin AI software engineering agent. The report covers the performanc

Parisians: we're running an open source AI art hackathon with LTX + NVIDIA this Saturday

HardwareDGX agent

A Reddit post on r/StableDiffusion announces an in-person open source AI art hackathon held in Paris, co-organized with LTX (Lightricks' open source AI video model) and NVIDIA. The event likely invite

Please Make it Sound like Human: Encoder-Decoder vs. Decoder-Only Transformers for AI-to-Human Text Style Transfer

Model ReleasesDGX agent

arXiv:2604.11687v1 Announce Type: new Abstract: AI-generated text has become common in academic and professional writing, prompting research into detection methods. Less studied is the reverse: system

PRISM Risk Signal Framework: Hierarchy-Based Red Lines for AI Behavioral Risk

SafetyDGX agent

arXiv:2604.11070v1 Announce Type: new Abstract: Current approaches to AI safety define red lines at the case level: specific prompts, specific outputs, specific harms. This paper argues that red lines

Progressive Multimodal Interaction Network for Reliable Quantification of Fish Feeding Intensity in Aquaculture

Model ReleasesDGX agent

arXiv:2506.14170v3 Announce Type: replace-cross Abstract: Accurate quantification of fish feeding intensity is crucial for precision feeding in aquaculture, as it directly affects feed utilization and

Quantization Dominates Rank Reduction for KV-Cache Compression

Model ReleasesDGX agent

arXiv:2604.11501v1 Announce Type: cross Abstract: We compare two strategies for compressing the KV cache in transformer inference: rank reduction (discard dimensions) and quantization (keep all dimens

Relational Preference Encoding in Looped Transformer Internal States

Model ReleasesDGX agent

arXiv:2604.09870v1 Announce Type: cross Abstract: We investigate how looped transformers encode human preference in their internal iteration states. Using Ouro-2.6B-Thinking, a 2.6B-parameter looped t

Robust Adversarial Policy Optimization Under Dynamics Uncertainty

Model ReleasesDGX agent

arXiv:2604.10974v1 Announce Type: new Abstract: Reinforcement learning (RL) policies often fail under dynamics that differ from training, a gap not fully addressed by domain randomization or existing

Saar-Voice: A Multi-Speaker Saarbrucken Dialect Speech Corpus

Local AiDGX agent

arXiv:2604.11803v1 Announce Type: new Abstract: Natural language processing (NLP) and speech technologies have made significant progress in recent years; however, they remain largely focused on standa

See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment

SafetyDGX agent

arXiv:2604.09749v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) frequently hallucinate objects that are absent from the visual input, often because attention during decoding i

Semantic Manipulation Localization

Model ReleasesDGX agent

arXiv:2604.10132v1 Announce Type: cross Abstract: Image Manipulation Localization (IML) aims to identify edited regions in an image. However, with the increasing use of modern image editing and genera

Seven simple steps for log analysis in AI systems

ResearchDGX agent

arXiv:2604.09563v1 Announce Type: new Abstract: AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensitie

Sheaf Diffusion with Adaptive Local Structure for Spatio-Temporal Forecasting

Local AiDGX agent

arXiv:2604.11275v1 Announce Type: new Abstract: Spatio-temporal systems often exhibit highly heterogeneous and non-intuitive responses to localized disruptions, limiting the effectiveness of conventio

Simulating Organized Group Behavior: New Framework, Benchmark, and Analysis

Model ReleasesDGX agent

arXiv:2604.09874v1 Announce Type: new Abstract: Simulating how organized groups (e.g., corporations) make decisions (e.g., responding to a competitor's move) is essential for understanding real-world

Simulator Adaptation for Sim-to-Real Learning of Legged Locomotion via Proprioceptive Distribution Matching

Model ReleasesDGX agent

arXiv:2604.11090v1 Announce Type: new Abstract: Simulation trained legged locomotion policies often exhibit performance loss on hardware due to dynamics discrepancies between the simulator and the rea

SLALOM: Simulation Lifecycle Analysis via Longitudinal Observation Metrics for Social Simulation

SafetyDGX agent

arXiv:2604.11466v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a potentially-transformative path forward for generative social science but face a critical crisis of validity

Spatial Competence Benchmark

Model ReleasesDGX agent

arXiv:2604.09594v1 Announce Type: new Abstract: Spatial competence is the quality of maintaining a consistent internal representation of an environment and using it to infer discrete structure and pla

SpotFormer: Multi-Scale Spatio-Temporal Transformer for Facial Expression Spotting

Local AiDGX agent

arXiv:2407.20799v3 Announce Type: replace Abstract: Facial expression spotting, identifying periods where facial expressions occur in a video, is a significant yet challenging task in facial expressio

← Previous
1…497498499500501…1071
Next →