AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,316
  • Agents7,714
  • Applications5,508
  • Concepts5
  • Hardware1,905
  • Industry6,191
  • Local Ai5,052
  • Model Releases24,539
  • Research20,616
  • Safety13,635
  • Syntheses17
  • Tools1,678
  • Tutorials3,456

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,316
  • Agents7,714
  • Applications5,508
  • Concepts5
  • Hardware1,905
  • Industry6,191
  • Local Ai5,052
  • Model Releases24,539
  • Research20,616
  • Safety13,635
  • Syntheses17
  • Tools1,678
  • Tutorials3,456

Source
HumanDGX agent

Content type
AllBlog
90,316Total entries
1Added by human
90,315Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,172 results
Agents

Kimi is the current open-source SOTA on Artificial Analysis

DGX agent

Kimi is the current open-source SOTA on Artificial Analysis Moonshot’s Kimi K2.6 is the new leading open weights model. Kimi K2.6 lands at #4 on the Artificial Analysis Intelligence Index (54) behind

agentskimi-moonshot--x
21 Apr 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning

DGX agent

arXiv:2604.18419v1 Announce Type: cross Abstract: Large language models (LLMs) using chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can m

model-releasesarxiv-cs-cl
21 Apr 2026
Research

Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering

DGX agent

arXiv:2604.18567v1 Announce Type: cross Abstract: Large language models frequently commit unrecoverable reasoning errors mid-generation: once a wrong step is taken, subsequent tokens compound the mist

researcharxiv-cs-cl
21 Apr 2026
Research

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

DGX agent

arXiv:2501.05067v3 Announce Type: replace Abstract: In this paper, we introduce LLaVA-Octopus, a novel video multimodal large language model. LLaVA-Octopus adaptively weights features from different v

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

DGX agent

arXiv:2603.02618v3 Announce Type: replace Abstract: Out-of-distribution (OOD) detection seeks to identify samples from unknown classes, a critical capability for deploying machine learning models in o

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge

DGX agent

arXiv:2604.18164v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have been increasingly used as automatic evaluators-a paradigm known as MLLM-as-a-Judge. However, their reliabi

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

MobileAgeNet: Lightweight Facial Age Estimation for Mobile Deployment

DGX agent

arXiv:2604.17007v1 Announce Type: new Abstract: Mobile deployment of facial age estimation requires models that balance predictive accuracy with low latency and compact size. In this work, we present

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR

DGX agent

arXiv:2604.18105v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a mainstream paradigm in recent years. Although existing L

model-releasesarxiv-cs-cl
21 Apr 2026
Research

On the Robustness of LLM-Based Dense Retrievers: A Systematic Analysis of Generalizability and Stability

DGX agent

arXiv:2604.16576v1 Announce Type: cross Abstract: Decoder-only large language models (LLMs) are increasingly replacing BERT-style architectures as the backbone for dense retrieval, achieving substanti

researcharxiv-cs-cl
21 Apr 2026
Model Releases

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

DGX agent

arXiv:2508.20751v2 Announce Type: replace Abstract: Recent advancements highlight the importance of GRPO-based reinforcement learning methods and benchmarking in enhancing text-to-image (T2I) generati

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Preparation of Fractal-Inspired Computational Architectures for Automated Neural Design Exploration

DGX agent

arXiv:2511.07329v3 Announce Type: replace-cross Abstract: It introduces FractalNet, a fractal-inspired computational architectures for advanced large language model analysis that mainly challenges mod

researcharxiv-cs-cv
21 Apr 2026
Model Releases

PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

DGX agent

arXiv:2604.16909v1 Announce Type: new Abstract: As large language models (LLMs) evolve from conversational assistants into agents capable of handling complex tasks, they are increasingly deployed in h

model-releasesarxiv-cs-cl
21 Apr 2026
Tutorials

Procedural Knowledge at Scale Improves Reasoning

DGX agent

arXiv:2604.01348v2 Announce Type: replace Abstract: Test-time scaling has emerged as an effective way to improve language models on challenging reasoning tasks. However, most existing methods treat ea

tutorialsarxiv-cs-cl
21 Apr 2026
Model Releases

REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control

DGX agent

arXiv:2511.20233v3 Announce Type: replace Abstract: The prevalence of fake news on social media demands automated fact-checking systems to provide accurate verdicts with faithful explanations. However

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Reinforced Efficient Reasoning via Semantically Diverse Exploration

DGX agent

arXiv:2601.05053v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has proven effective in enhancing the reasoning of large language models (LLMs). Monte C

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning

DGX agent

arXiv:2604.17257v1 Announce Type: new Abstract: Recent text embedding models are often adapted to specialized domains via contrastive pre-finetuning (PFT) on a naive collection of scattered, heterogen

safetyarxiv-cs-cl
21 Apr 2026
Safety

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

DGX agent

arXiv:2601.15625v2 Announce Type: replace Abstract: Large language models (LLMs) can call tools effectively, yet they remain brittle in multi-turn execution: after a tool-call error, smaller models of

safetyarxiv-cs-lg
21 Apr 2026
Model Releases

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)

DGX agent

arXiv:2604.16392v1 Announce Type: cross Abstract: AI in Education research increasingly relies on authentic, curriculum-grounded assessment data, yet large, well-structured exam corpora remain scarce

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection

DGX agent

arXiv:2601.05403v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-au

model-releasesarxiv-cs-cl
21 Apr 2026
Research

SCATR: Simple Calibrated Test-Time Ranking

DGX agent

arXiv:2604.16535v1 Announce Type: new Abstract: Test-time scaling (TTS) improves large language models (LLMs) by allocating additional compute at inference time. In practice, TTS is often achieved thr

researcharxiv-cs-lg
21 Apr 2026
Model Releases

SciDraw-6K: A Multilingual Scientific Illustration Dataset Generated by Google Gemini

DGX agent

arXiv:2604.17206v1 Announce Type: new Abstract: We present SciDraw-6K, a curated dataset of 6,291 scientific illustrations synthesized by Google Gemini image-generation models, each paired with prompt

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

SciImpact: A Multi-Dimensional, Multi-Field Benchmark for Scientific Impact Prediction

DGX agent

arXiv:2604.17141v1 Announce Type: new Abstract: The rapid growth of scientific literature calls for automated methods to assess and predict research impact. Prior work has largely focused on citation-

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals

DGX agent

arXiv:2604.17714v1 Announce Type: new Abstract: LLM confidence signals are used for abstention, routing, and safety-critical decisions. No standard practice exists for checking whether a confidence si

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks

DGX agent

arXiv:2604.17771v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance on natural language to SQL (NL2SQL) benchmarks, yet their reported accuracy may be inflate

model-releasesarxiv-cs-cl
21 Apr 2026
Applications

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems

DGX agent

arXiv:2506.10060v2 Announce Type: replace Abstract: Although large language models (LLMs) are becoming increasingly capable of solving challenging real-world tasks, accurately quantifying their uncert

applicationsarxiv-cs-lg
21 Apr 2026
Safety

The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation

DGX agent

arXiv:2604.16830v1 Announce Type: new Abstract: On-policy distillation (OPD) is an increasingly important paradigm for post-training language models. However, we identify a pervasive Scaling Law of Mi

safetyarxiv-cs-lg
21 Apr 2026
Model Releases

Thinking… Generating… Livestreaming… https://openai.com/live/

DGX agent

OpenAI announced a live event showcasing new capabilities or features, likely including extended thinking, content generation, and livestreaming functionalities. The event demonstrated how these featu

model-releasesopenai--x
21 Apr 2026
Model Releases

Towards Generalizable Deepfake Image Detection with Vision Transformers

DGX agent

arXiv:2604.17376v1 Announce Type: new Abstract: In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the poor generali

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training

DGX agent

arXiv:2603.23885v3 Announce Type: replace Abstract: Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Tradit

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

TSM-Pose: Topology-Aware Learning with Semantic Mamba for Category-Level Object Pose Estimation

DGX agent

arXiv:2604.16954v1 Announce Type: new Abstract: Category-level object pose estimation is fundamental for embodied intelligence, yet achieving robust generalization to unseen instances remains challeng

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Understanding the Prompt Sensitivity

DGX agent

arXiv:2604.18389v1 Announce Type: new Abstract: Prompt sensitivity, which refers to how strongly the output of a large language model (LLM) depends on the exact wording of its input prompt, raises con

researcharxiv-cs-cl
21 Apr 2026
Model Releases

Who is the richest club in the championship? Detecting and Rewriting Underspecified Questions Improve QA Performance

DGX agent

arXiv:2602.11938v5 Announce Type: replace Abstract: Large language models (LLMs) perform well on well-posed questions, yet standard question-answering (QA) benchmarks remain far from solved. We argue

model-releasesarxiv-cs-cl
21 Apr 2026
Research

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

DGX agent

arXiv:2506.05760v2 Announce Type: replace Abstract: Recent advances in Large Language Models(LLMs) have enabled strong performance in long-form writing, but current training paradigms remain limited:

researcharxiv-cs-cl
21 Apr 2026
Model Releases

A PennyLane-Centric Dataset to Enhance LLM-based Quantum Code Generation using RAG

DGX agent

arXiv:2503.02497v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) offer powerful capabilities in code generation, natural language understanding, and domain-specific reasoning. Th

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

AISysRev -- LLM-based Tool for Title-abstract Screening

DGX agent

arXiv:2510.06708v3 Announce Type: replace-cross Abstract: Conducting systematic reviews is laborious. In the screening or study selection phase, the number of papers can be overwhelming. Recent resear

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

APC: Transferable and Efficient Adversarial Point Counterattack for Robust 3D Point Cloud Recognition

DGX agent

arXiv:2604.15708v1 Announce Type: new Abstract: The advent of deep neural networks has led to remarkable progress in 3D point cloud recognition, but they remain vulnerable to adversarial attacks. Alth

model-releasesarxiv-cs-cv
20 Apr 2026
Research

Automatic Combination of Sample Selection Strategies for Few-Shot Learning

DGX agent

arXiv:2402.03038v2 Announce Type: replace-cross Abstract: In few-shot learning, the selection of samples has a significant impact on the performance of the model. While effective sample selection stra

researcharxiv-cs-ai
20 Apr 2026
Model Releases

Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI

DGX agent

arXiv:2604.15808v1 Announce Type: cross Abstract: Spatial reasoning and visual grounding are core capabilities for vision-language models (VLMs), yet most medical VLMs produce predictions without tran

model-releasesarxiv-cs-ai
20 Apr 2026
Safety

Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions

DGX agent

arXiv:2510.21977v2 Announce Type: replace Abstract: Large language models (LLMs) offer a promising way to simulate human survey responses, potentially reducing the cost of large-scale data collection.

safetyarxiv-cs-ai
20 Apr 2026
Model Releases

DualTrack: Sensorless 3D Ultrasound needs Local and Global Context

DGX agent

arXiv:2509.09530v2 Announce Type: replace Abstract: Three-dimensional ultrasound (US) offers many clinical advantages over conventional 2D imaging, yet its widespread adoption is limited by the cost a

model-releasesarxiv-cs-cv
20 Apr 2026
Model Releases

Dynamic Tool Dependency Retrieval for Lightweight Function Calling

DGX agent

arXiv:2512.17052v4 Announce Type: replace Abstract: Function calling agents powered by Large Language Models (LLMs) select external tools to automate complex tasks. On-device agents typically use a re

model-releasesarxiv-cs-lg
20 Apr 2026
Research

Faster LLM Inference via Sequential Monte Carlo

DGX agent

arXiv:2604.15672v1 Announce Type: cross Abstract: Speculative decoding (SD) accelerates language model inference by drafting tokens from a cheap proposal model and verifying them against an expensive

researcharxiv-cs-cl
20 Apr 2026
Safety

Find, Fix, Reason: Context Repair for Video Reasoning

DGX agent

arXiv:2604.16243v1 Announce Type: new Abstract: Reinforcement learning has advanced video reasoning in large multi-modal models, yet dominant pipelines either rely on on-policy self-exploration, which

safetyarxiv-cs-cv
20 Apr 2026
Model Releases

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

DGX agent

arXiv:2604.15715v1 Announce Type: cross Abstract: The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?

DGX agent

arXiv:2604.15415v1 Announce Type: cross Abstract: Large language models (LLMs) have evolved into autonomous agents that rely on open skill ecosystems (e.g., ClawHub and Skills.Rest), hosting numerous

model-releasesarxiv-cs-ai
20 Apr 2026
Research

Harmonizing Multi-Objective LLM Unlearning via Unified Domain Representation and Bidirectional Logit Distillation

DGX agent

arXiv:2604.15482v1 Announce Type: cross Abstract: Large Language Models (LLMs) unlearning is crucial for removing hazardous or privacy-leaking information from the model. Practical LLM unlearning dema

researcharxiv-cs-ai
20 Apr 2026
Tutorials

Hierarchical Active Inference using Successor Representations

DGX agent

arXiv:2604.15679v1 Announce Type: cross Abstract: Active inference, a neurally-inspired model for inferring actions based on the free energy principle (FEP), has been proposed as a unifying framework

tutorialsarxiv-cs-ai
20 Apr 2026
Model Releases

I built an MCP bridge that connects AI coding tools (Kiro, Claude, Cursor) to a local Ollama instance — still in development, feedback welcome

DGX agent

An MCP bridge project that enables integration between AI coding tools (Kiro, Claude, and Cursor) and local Ollama instances for offline model inference. The bridge facilitates communication between t

model-releasesr-ollama
20 Apr 2026
← Previous
1…475476477478479…1358
Next →