AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,315 results
21 Apr 2026

Predicting LLM Compression Degradation from Spectral Statistics

Model ReleasesDGX agent

arXiv:2604.18085v1 Announce Type: new Abstract: Matrix-level low-rank compression is a promising way to reduce the cost of large language models, but running compression and evaluating the resulting m

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

Model ReleasesDGX agent

arXiv:2508.20751v2 Announce Type: replace Abstract: Recent advancements highlight the importance of GRPO-based reinforcement learning methods and benchmarking in enhancing text-to-image (T2I) generati

PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2506.13674v3 Announce Type: replace Abstract: Parameter-Efficient Fine-Tuning (PEFT) methods have become crucial for rapidly adapting large language models (LLMs) to downstream tasks. Prefix-Tun

Prior-Fitted Functional Flow: In-Context Generative Models for Pharmacokinetics

Model ReleasesDGX agent

arXiv:2604.17670v1 Announce Type: new Abstract: We introduce Prior-Fitted Functional Flows, a generative foundation model for pharmacokinetics that enables zero-shot population synthesis and individua

PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

Model ReleasesDGX agent

arXiv:2604.16909v1 Announce Type: new Abstract: As large language models (LLMs) evolve from conversational assistants into agents capable of handling complex tasks, they are increasingly deployed in h

ProfVLM: A lightweight video-language model for multi-view proficiency estimation

Model ReleasesDGX agent

arXiv:2509.26278v4 Announce Type: replace-cross Abstract: Most existing approaches formulate action quality assessment and skill proficiency estimation as discriminative prediction tasks, typically pr

Program Structure-aware Language Models: Targeted Software Testing beyond Textual Semantics

Model ReleasesDGX agent

arXiv:2604.17715v1 Announce Type: cross Abstract: Recent advances in large language models for test case generation have improved branch coverage via prompt-engineered mutations. However, they still l

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

Model ReleasesDGX agent

arXiv:2604.18459v1 Announce Type: new Abstract: Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability th

Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery

Model ReleasesDGX agent

arXiv:2604.17920v1 Announce Type: new Abstract: Synthetic Aperture Radar (SAR) plays a critical role in maritime surveillance, yet deep learning for SAR analysis is limited by the lack of pixel-level

ProTrain: Efficient LLM Training via Memory-Aware Techniques

Model ReleasesDGX agent

arXiv:2406.08334v2 Announce Type: replace-cross Abstract: Memory pressure has emerged as a dominant constraint in scaling the training of large language models (LLMs), particularly in resource-constra

Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution

Model ReleasesDGX agent

arXiv:2604.16889v1 Announce Type: new Abstract: Existing feature-interpretation pipelines typically operate on uniformly sampled units, but only a small fraction of cross-layer transcoder (CLT) featur

Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition

Model ReleasesDGX agent

arXiv:2510.08047v2 Announce Type: replace-cross Abstract: Robust ASR under domain shift is crucial because real-world systems encounter unseen accents and domains with limited labeled data. Although p

PTLD: Sim-to-real Privileged Tactile Latent Distillation for Dexterous Manipulation

Model ReleasesDGX agent

arXiv:2603.04531v2 Announce Type: replace Abstract: Tactile dexterous manipulation is essential to automating complex household tasks, yet learning effective control policies remains a challenge. Whil

Pulse Shape Discrimination Algorithms: Survey and Benchmark

Model ReleasesDGX agent

arXiv:2508.02750v2 Announce Type: replace Abstract: This review presents a comprehensive survey and benchmark of pulse shape discrimination (PSD) algorithms for radiation detection, classifying nearly

QU-NLP at QIAS 2026: Multi-Stage QLoRA Fine-Tuning for Arabic Islamic Inheritance Reasoning

Model ReleasesDGX agent

arXiv:2604.16396v1 Announce Type: new Abstract: Islamic inheritance law (ilm al-mawar{i}th) presents a challenging domain for evaluating large language models' structured reasoning capabilities, requi

ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval

Model ReleasesDGX agent

arXiv:2510.08252v2 Announce Type: replace-cross Abstract: In this paper, we introduce ReasonEmbed, a novel text embedding model developed for reasoning-intensive document retrieval. Our work includes

ReCap: Lightweight Referential Grounding for Coherent Story Visualization

Model ReleasesDGX agent

arXiv:2604.18575v1 Announce Type: new Abstract: Story Visualization aims to generate a sequence of images that faithfully depicts a textual narrative that preserve character identity, spatial configur

ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering

Model ReleasesDGX agent

arXiv:2604.17944v1 Announce Type: new Abstract: Developing agents capable of navigating fragmented, multi-source information remains challenging, primarily due to the scarcity of benchmarks reflecting

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning

Model ReleasesDGX agent

arXiv:2604.17800v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have gained much attention from the research community thanks to their strength in translating multimodal observat

REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control

Model ReleasesDGX agent

arXiv:2511.20233v3 Announce Type: replace Abstract: The prevalence of fake news on social media demands automated fact-checking systems to provide accurate verdicts with faithful explanations. However

ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2603.05863v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have revolutionized code generation, standard ``System 1'' approaches that generate solutions in a single forward

Reinforced Efficient Reasoning via Semantically Diverse Exploration

Model ReleasesDGX agent

arXiv:2601.05053v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has proven effective in enhancing the reasoning of large language models (LLMs). Monte C

Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning

Model ReleasesDGX agent

arXiv:2601.02970v2 Announce Type: replace Abstract: Self-Consistency improves reasoning reliability through multi-sample aggregation, but incurs substantial inference cost. Adaptive self-consistency m

Representation Before Training: A Fixed-Budget Benchmark for Generative Medical Event Models

Model ReleasesDGX agent

arXiv:2604.16775v1 Announce Type: new Abstract: Every prediction from a generative medical event model is bounded by how clinical events are tokenized, yet input representation is rarely isolated from

Representation-Guided Parameter-Efficient LLM Unlearning

Model ReleasesDGX agent

arXiv:2604.17396v1 Announce Type: new Abstract: Large Language Models (LLMs) often memorize sensitive or harmful information, necessitating effective machine unlearning techniques. While existing para

ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Model ReleasesDGX agent

arXiv:2503.21248v3 Announce Type: replace Abstract: Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses r

Rethinking Cross-Modal Fine-Tuning: Optimizing the Interaction Between Feature Alignment and Target Fitting

Model ReleasesDGX agent

arXiv:2601.18231v4 Announce Type: replace Abstract: Adapting pre-trained models to unseen feature modalities has become increasingly important due to the growing need for cross-disciplinary knowledge

Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation

Model ReleasesDGX agent

arXiv:2604.17260v1 Announce Type: new Abstract: Evaluating meeting effectiveness is crucial for improving organizational productivity. Current approaches rely on post-hoc surveys that yield a single c

ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering

Model ReleasesDGX agent

arXiv:2510.09351v2 Announce Type: replace Abstract: While Small Language Models (SLMs) have demonstrated promising performance on an increasingly wide array of commonsense reasoning benchmarks, curren

ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval

Model ReleasesDGX agent

arXiv:2604.17898v1 Announce Type: new Abstract: With the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing atten

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models

Model ReleasesDGX agent

arXiv:2604.16593v1 Announce Type: new Abstract: We present SemanticQA, an evaluation suite designed to assess language models (LMs) in semantic phrase processing tasks. The benchmark consolidates exis

Revisiting Active Sequential Prediction-Powered Mean Estimation

Model ReleasesDGX agent

arXiv:2604.18569v1 Announce Type: cross Abstract: In this work, we revisit the problem of active sequential prediction-powered mean estimation, where at each round one must decide the query probabilit

Revisiting Auxiliary Losses for Conditional Depth Routing: An Empirical Study

Model ReleasesDGX agent

arXiv:2604.17228v1 Announce Type: new Abstract: Conditional depth execution routes a subset of tokens through a lightweight cheap FFN while the remainder execute the standard full FFN at each controll

Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models

Model ReleasesDGX agent

arXiv:2604.18429v1 Announce Type: new Abstract: Change visual question answering (Change VQA) addresses the problem of answering natural-language questions about semantic changes between bi-temporal r

R&F-Inventory: A Large-Scale Dataset for Monotonic Inventory Estimation in Reach and Frequency Advertising

Model ReleasesDGX agent

arXiv:2604.16821v1 Announce Type: new Abstract: Reach and Frequency (R&F) contract advertising is an important form of widely used brand advertising. Unlike performance advertising, R&F contracts emph

Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models

Model ReleasesDGX agent

arXiv:2602.14466v2 Announce Type: replace Abstract: With natural language generation becoming a popular use case for language models, the Bias Benchmark for Question-Answering (BBQ) has grown to be an

RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian

Model ReleasesDGX agent

arXiv:2604.17134v1 Announce Type: new Abstract: We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)

Model ReleasesDGX agent

arXiv:2604.16392v1 Announce Type: cross Abstract: AI in Education research increasingly relies on authentic, curriculum-grounded assessment data, yet large, well-structured exam corpora remain scarce

RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design

Model ReleasesDGX agent

arXiv:2604.17175v1 Announce Type: new Abstract: We introduce RosettaSearch, an inference-time multi-objective optimization approach for protein sequence optimization. We use large language models (LLM

RoTRAG: Rule of Thumb Reasoning for Conversation Harm Detection with Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2604.17301v1 Announce Type: new Abstract: Detecting harmful content in multi turn dialogue requires reasoning over the full conversational context rather than isolated utterances. However, most

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification

Model ReleasesDGX agent

arXiv:2604.16993v1 Announce Type: cross Abstract: As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachabili

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.18512v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated remarkable progress in single-image understanding, yet effective reasoning across multiple images remain

Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape

Model ReleasesDGX agent

arXiv:2505.21722v2 Announce Type: replace Abstract: When a deep ReLU network is initialized with small weights, gradient descent (GD) is at first dominated by the saddle at the origin in parameter spa

Safe Control using Learned Safety Filters and Adaptive Conformal Inference

Model ReleasesDGX agent

arXiv:2604.18482v1 Announce Type: cross Abstract: Safety filters have been shown to be effective tools to ensure the safety of control systems with unsafe nominal policies. To address scalability chal

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models

Model ReleasesDGX agent

arXiv:2604.17691v1 Announce Type: new Abstract: Safety alignment in large language models is remarkably shallow: it is concentrated in the first few output tokens and reversible by fine-tuning on as f

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning

Model ReleasesDGX agent

arXiv:2503.03480v4 Announce Type: replace Abstract: Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-w

Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection

Model ReleasesDGX agent

arXiv:2601.05403v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-au

Sampling Matters: The Effect of ECG Frequency on Deep Learning-Based Atrial Fibrillation Detection

Model ReleasesDGX agent

arXiv:2604.16437v1 Announce Type: cross Abstract: Deep learning models for atrial fibrillation (AF) detection are increasingly trained on heterogeneous electrocardiogram (ECG) datasets with varying sa

SAUVER LA FRANCE ET L'EUROPE EN FAISANT LE SEUL PARI QUI VAILLE (série banane rouge) I. COMME BONAPARTE : CONCENTRER LES FORCES, PAS LES DIL…

Model ReleasesDGX agent

SAUVER LA FRANCE ET L'EUROPE EN FAISANT LE SEUL PARI QUI VAILLE (série banane rouge) I. COMME BONAPARTE : CONCENTRER LES FORCES, PAS LES DILUER Dans le discours politico-industriel d'aujourd'hui, on e

Scaling Codex to enterprises worldwide

Model ReleasesDGX agent

OpenAI's Codex, a code generation model trained on publicly available code, was scaled for enterprise deployment to provide developers with AI-assisted coding capabilities across organizations worldwi

Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration

Model ReleasesDGX agent

arXiv:2505.21471v2 Announce Type: replace Abstract: With the rapid advancement of post-training techniques for reasoning and information seeking, large language models (LLMs) can incorporate a large q

Scaling Human-AI Coding Collaboration Requires a Governable Consensus Layer

Model ReleasesDGX agent

arXiv:2604.17883v1 Announce Type: cross Abstract: Vibe coding produces correct, executable code at speed, but leaves no record of the structural commitments, dependencies, or evidence behind it. Revie

Scaling Recurrence-aware Foundation Models for Clinical Records via Next-Visit Prediction

Model ReleasesDGX agent

arXiv:2603.24562v2 Announce Type: replace Abstract: While large-scale pretraining has revolutionized language modeling, its potential remains underexplored in healthcare with structured electronic hea

Scaling Test-Time Compute for Agentic Coding

Model ReleasesDGX agent

arXiv:2604.16529v1 Announce Type: cross Abstract: Test-time scaling has become a powerful way to improve large language models. However, existing methods are best suited to short, bounded outputs that

Scaling up Kimi K2.6 with more NVIDIA Blackwell GPUs! Will be adding more in the coming days. Try it with OpenClaw: ollama launch openclaw -…

Model ReleasesDGX agent

Scaling up Kimi K2.6 with more NVIDIA Blackwell GPUs! Will be adding more in the coming days. Try it with OpenClaw: ollama launch openclaw --model kimi-k2.6:cloud Try it with Hermes Agent: ollama laun

SciDraw-6K: A Multilingual Scientific Illustration Dataset Generated by Google Gemini

Model ReleasesDGX agent

arXiv:2604.17206v1 Announce Type: new Abstract: We present SciDraw-6K, a curated dataset of 6,291 scientific illustrations synthesized by Google Gemini image-generation models, each paired with prompt

ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

Model ReleasesDGX agent

arXiv:2505.19897v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have extended their impact beyond Natural Language Processing, substantially fostering the development of interdi

SciImpact: A Multi-Dimensional, Multi-Field Benchmark for Scientific Impact Prediction

Model ReleasesDGX agent

arXiv:2604.17141v1 Announce Type: new Abstract: The rapid growth of scientific literature calls for automated methods to assess and predict research impact. Prior work has largely focused on citation-

Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals

Model ReleasesDGX agent

arXiv:2604.17714v1 Announce Type: new Abstract: LLM confidence signals are used for abstention, routing, and safety-critical decisions. No standard practice exists for checking whether a confidence si

SeekerGym: A Benchmark for Reliable Information Seeking

Model ReleasesDGX agent

arXiv:2604.17143v1 Announce Type: new Abstract: Despite their substantial successes, AI agents continue to face fundamental challenges in terms of trustworthiness. Consider deep research agents, taske

← Previous
1…329330331332333…372
Next →