AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,340 results
Model Releases

Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?

DGX agent

arXiv:2604.17338v1 Announce Type: cross Abstract: Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but o

model-releasesarxiv-cs-cl
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Predicting LLM Compression Degradation from Spectral Statistics

DGX agent

arXiv:2604.18085v1 Announce Type: new Abstract: Matrix-level low-rank compression is a promising way to reduce the cost of large language models, but running compression and evaluating the resulting m

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

DGX agent

arXiv:2508.20751v2 Announce Type: replace Abstract: Recent advancements highlight the importance of GRPO-based reinforcement learning methods and benchmarking in enhancing text-to-image (T2I) generati

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention

DGX agent

arXiv:2506.13674v3 Announce Type: replace Abstract: Parameter-Efficient Fine-Tuning (PEFT) methods have become crucial for rapidly adapting large language models (LLMs) to downstream tasks. Prefix-Tun

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Prior-Fitted Functional Flow: In-Context Generative Models for Pharmacokinetics

DGX agent

arXiv:2604.17670v1 Announce Type: new Abstract: We introduce Prior-Fitted Functional Flows, a generative foundation model for pharmacokinetics that enables zero-shot population synthesis and individua

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

DGX agent

arXiv:2604.16909v1 Announce Type: new Abstract: As large language models (LLMs) evolve from conversational assistants into agents capable of handling complex tasks, they are increasingly deployed in h

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ProfVLM: A lightweight video-language model for multi-view proficiency estimation

DGX agent

arXiv:2509.26278v4 Announce Type: replace-cross Abstract: Most existing approaches formulate action quality assessment and skill proficiency estimation as discriminative prediction tasks, typically pr

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Program Structure-aware Language Models: Targeted Software Testing beyond Textual Semantics

DGX agent

arXiv:2604.17715v1 Announce Type: cross Abstract: Recent advances in large language models for test case generation have improved branch coverage via prompt-engineered mutations. However, they still l

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

DGX agent

arXiv:2604.18459v1 Announce Type: new Abstract: Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability th

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery

DGX agent

arXiv:2604.17920v1 Announce Type: new Abstract: Synthetic Aperture Radar (SAR) plays a critical role in maritime surveillance, yet deep learning for SAR analysis is limited by the lack of pixel-level

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

ProTrain: Efficient LLM Training via Memory-Aware Techniques

DGX agent

arXiv:2406.08334v2 Announce Type: replace-cross Abstract: Memory pressure has emerged as a dominant constraint in scaling the training of large language models (LLMs), particularly in resource-constra

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution

DGX agent

arXiv:2604.16889v1 Announce Type: new Abstract: Existing feature-interpretation pipelines typically operate on uniformly sampled units, but only a small fraction of cross-layer transcoder (CLT) featur

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition

DGX agent

arXiv:2510.08047v2 Announce Type: replace-cross Abstract: Robust ASR under domain shift is crucial because real-world systems encounter unseen accents and domains with limited labeled data. Although p

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

PTLD: Sim-to-real Privileged Tactile Latent Distillation for Dexterous Manipulation

DGX agent

arXiv:2603.04531v2 Announce Type: replace Abstract: Tactile dexterous manipulation is essential to automating complex household tasks, yet learning effective control policies remains a challenge. Whil

model-releasesarxiv-cs-ro
21 Apr 2026
Model Releases

Pulse Shape Discrimination Algorithms: Survey and Benchmark

DGX agent

arXiv:2508.02750v2 Announce Type: replace Abstract: This review presents a comprehensive survey and benchmark of pulse shape discrimination (PSD) algorithms for radiation detection, classifying nearly

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

QU-NLP at QIAS 2026: Multi-Stage QLoRA Fine-Tuning for Arabic Islamic Inheritance Reasoning

DGX agent

arXiv:2604.16396v1 Announce Type: new Abstract: Islamic inheritance law (ilm al-mawar{i}th) presents a challenging domain for evaluating large language models' structured reasoning capabilities, requi

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval

DGX agent

arXiv:2510.08252v2 Announce Type: replace-cross Abstract: In this paper, we introduce ReasonEmbed, a novel text embedding model developed for reasoning-intensive document retrieval. Our work includes

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ReCap: Lightweight Referential Grounding for Coherent Story Visualization

DGX agent

arXiv:2604.18575v1 Announce Type: new Abstract: Story Visualization aims to generate a sequence of images that faithfully depicts a textual narrative that preserve character identity, spatial configur

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering

DGX agent

arXiv:2604.17944v1 Announce Type: new Abstract: Developing agents capable of navigating fragmented, multi-source information remains challenging, primarily due to the scarcity of benchmarks reflecting

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning

DGX agent

arXiv:2604.17800v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have gained much attention from the research community thanks to their strength in translating multimodal observat

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control

DGX agent

arXiv:2511.20233v3 Announce Type: replace Abstract: The prevalence of fake news on social media demands automated fact-checking systems to provide accurate verdicts with faithful explanations. However

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning

DGX agent

arXiv:2603.05863v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have revolutionized code generation, standard ``System 1'' approaches that generate solutions in a single forward

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Reinforced Efficient Reasoning via Semantically Diverse Exploration

DGX agent

arXiv:2601.05053v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has proven effective in enhancing the reasoning of large language models (LLMs). Monte C

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning

DGX agent

arXiv:2601.02970v2 Announce Type: replace Abstract: Self-Consistency improves reasoning reliability through multi-sample aggregation, but incurs substantial inference cost. Adaptive self-consistency m

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Representation Before Training: A Fixed-Budget Benchmark for Generative Medical Event Models

DGX agent

arXiv:2604.16775v1 Announce Type: new Abstract: Every prediction from a generative medical event model is bounded by how clinical events are tokenized, yet input representation is rarely isolated from

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Representation-Guided Parameter-Efficient LLM Unlearning

DGX agent

arXiv:2604.17396v1 Announce Type: new Abstract: Large Language Models (LLMs) often memorize sensitive or harmful information, necessitating effective machine unlearning techniques. While existing para

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

DGX agent

arXiv:2503.21248v3 Announce Type: replace Abstract: Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses r

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Rethinking Cross-Modal Fine-Tuning: Optimizing the Interaction Between Feature Alignment and Target Fitting

DGX agent

arXiv:2601.18231v4 Announce Type: replace Abstract: Adapting pre-trained models to unseen feature modalities has become increasingly important due to the growing need for cross-disciplinary knowledge

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation

DGX agent

arXiv:2604.17260v1 Announce Type: new Abstract: Evaluating meeting effectiveness is crucial for improving organizational productivity. Current approaches rely on post-hoc surveys that yield a single c

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering

DGX agent

arXiv:2510.09351v2 Announce Type: replace Abstract: While Small Language Models (SLMs) have demonstrated promising performance on an increasingly wide array of commonsense reasoning benchmarks, curren

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval

DGX agent

arXiv:2604.17898v1 Announce Type: new Abstract: With the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing atten

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models

DGX agent

arXiv:2604.16593v1 Announce Type: new Abstract: We present SemanticQA, an evaluation suite designed to assess language models (LMs) in semantic phrase processing tasks. The benchmark consolidates exis

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Revisiting Active Sequential Prediction-Powered Mean Estimation

DGX agent

arXiv:2604.18569v1 Announce Type: cross Abstract: In this work, we revisit the problem of active sequential prediction-powered mean estimation, where at each round one must decide the query probabilit

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Revisiting Auxiliary Losses for Conditional Depth Routing: An Empirical Study

DGX agent

arXiv:2604.17228v1 Announce Type: new Abstract: Conditional depth execution routes a subset of tokens through a lightweight cheap FFN while the remainder execute the standard full FFN at each controll

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models

DGX agent

arXiv:2604.18429v1 Announce Type: new Abstract: Change visual question answering (Change VQA) addresses the problem of answering natural-language questions about semantic changes between bi-temporal r

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

R&F-Inventory: A Large-Scale Dataset for Monotonic Inventory Estimation in Reach and Frequency Advertising

DGX agent

arXiv:2604.16821v1 Announce Type: new Abstract: Reach and Frequency (R&F) contract advertising is an important form of widely used brand advertising. Unlike performance advertising, R&F contracts emph

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models

DGX agent

arXiv:2602.14466v2 Announce Type: replace Abstract: With natural language generation becoming a popular use case for language models, the Bias Benchmark for Question-Answering (BBQ) has grown to be an

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian

DGX agent

arXiv:2604.17134v1 Announce Type: new Abstract: We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)

DGX agent

arXiv:2604.16392v1 Announce Type: cross Abstract: AI in Education research increasingly relies on authentic, curriculum-grounded assessment data, yet large, well-structured exam corpora remain scarce

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design

DGX agent

arXiv:2604.17175v1 Announce Type: new Abstract: We introduce RosettaSearch, an inference-time multi-objective optimization approach for protein sequence optimization. We use large language models (LLM

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

RoTRAG: Rule of Thumb Reasoning for Conversation Harm Detection with Retrieval-Augmented Generation

DGX agent

arXiv:2604.17301v1 Announce Type: new Abstract: Detecting harmful content in multi turn dialogue requires reasoning over the full conversational context rather than isolated utterances. However, most

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification

DGX agent

arXiv:2604.16993v1 Announce Type: cross Abstract: As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachabili

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models

DGX agent

arXiv:2604.18512v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated remarkable progress in single-image understanding, yet effective reasoning across multiple images remain

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape

DGX agent

arXiv:2505.21722v2 Announce Type: replace Abstract: When a deep ReLU network is initialized with small weights, gradient descent (GD) is at first dominated by the saddle at the origin in parameter spa

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Safe Control using Learned Safety Filters and Adaptive Conformal Inference

DGX agent

arXiv:2604.18482v1 Announce Type: cross Abstract: Safety filters have been shown to be effective tools to ensure the safety of control systems with unsafe nominal policies. To address scalability chal

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models

DGX agent

arXiv:2604.17691v1 Announce Type: new Abstract: Safety alignment in large language models is remarkably shallow: it is concentrated in the first few output tokens and reversible by fine-tuning on as f

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning

DGX agent

arXiv:2503.03480v4 Announce Type: replace Abstract: Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-w

model-releasesarxiv-cs-ro
21 Apr 2026
Model Releases

Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection

DGX agent

arXiv:2601.05403v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-au

model-releasesarxiv-cs-cl
21 Apr 2026
← Previous
1…413414415416417…466
Next →