AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,106 results
Model Releases

Beyond Accuracy: A Stability-Aware Metric for Multi-Horizon Forecasting

DGX agent

arXiv:2601.10863v3 Announce Type: replace Abstract: Traditional time series forecasting methods optimize for accuracy alone. This objective neglects temporal consistency, in other words, how consisten

model-releasesarxiv-cs-lg
24 Apr 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Can MLLMs 'Read' What is Missing?

DGX agent

arXiv:2604.21277v1 Announce Type: new Abstract: We introduce MMTR-Bench, a benchmark designed to evaluate the intrinsic ability of Multimodal Large Language Models (MLLMs) to reconstruct masked text d

model-releasesarxiv-cs-ai
24 Apr 2026
Research

Certified Coil Geometry Learning for Short-Range Magnetic Actuation and Spacecraft Docking Application

DGX agent

arXiv:2507.03806v3 Announce Type: replace-cross Abstract: This paper presents a learning-based framework for approximating an exact magnetic-field interaction model, supported by both numerical and ex

researcharxiv-cs-lg
24 Apr 2026
Safety

Continuous-Utility Direct Preference Optimization

DGX agent

arXiv:2602.00931v2 Announce Type: replace-cross Abstract: Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture par

safetyarxiv-cs-ai
24 Apr 2026
Model Releases

DAVIS: OOD Detection via Dominant Activations and Variance for Increased Separation

DGX agent

arXiv:2601.22703v2 Announce Type: replace Abstract: Detecting out-of-distribution (OOD) inputs is a critical safeguard for deploying machine learning models in the real world. However, most post-hoc d

model-releasesarxiv-cs-cv
24 Apr 2026
Model Releases

Empirical Comparison of Agent Communication Protocols for Task Orchestration

DGX agent

arXiv:2603.22823v3 Announce Type: replace Abstract: Context. The problem of comparative evaluation of communication protocols for task orchestration by large language model (LLM) agents is considered.

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

EVENT5Ws: A Large Dataset for Open-Domain Event Extraction from Documents

DGX agent

arXiv:2604.21890v1 Announce Type: new Abstract: Event extraction identifies the central aspects of events from text. It supports event understanding and analysis, which is crucial for tasks such as in

model-releasesarxiv-cs-cl
24 Apr 2026
Research

Finding Meaning in Embeddings: Concept Separation Curves

DGX agent

arXiv:2604.21555v1 Announce Type: new Abstract: Sentence embedding techniques aim to encode key concepts of a sentence's meaning in a vector space. However, the majority of evaluation approaches for s

researcharxiv-cs-cl
24 Apr 2026
Research

Frequency-Forcing: From Scaling-as-Time to Soft Frequency Guidance

DGX agent

arXiv:2604.20902v1 Announce Type: cross Abstract: While standard flow-matching models transport noise to data uniformly, incorporating an explicit generation order - specifically, establishing coarse,

researcharxiv-cs-ai
24 Apr 2026
Model Releases

From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media

DGX agent

arXiv:2604.21786v1 Announce Type: new Abstract: Social media platforms have become primary arenas for climate communication, generating millions of images and posts that - if systematically analysed -

model-releasesarxiv-cs-cv
24 Apr 2026
Model Releases

Generalizing Numerical Reasoning in Table Data through Operation Sketches and Self-Supervised Learning

DGX agent

arXiv:2604.21495v1 Announce Type: cross Abstract: Numerical reasoning over expert-domain tables often exhibits high in-domain accuracy but limited robustness to domain shift. Models trained with super

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Geometric Characterisation and Structured Trajectory Surrogates for Clinical Dataset Condensation

DGX agent

arXiv:2604.21638v1 Announce Type: new Abstract: Dataset condensation constructs compact synthetic datasets that retain the training utility of large real-world datasets, enabling efficient model devel

model-releasesarxiv-cs-lg
24 Apr 2026
Model Releases

HyperAdapt: Simple High-Rank Adaptation

DGX agent

arXiv:2509.18629v3 Announce Type: replace-cross Abstract: Foundation models excel across diverse tasks, but adapting them to specialized applications often requires fine-tuning, an approach that is me

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Ideological Bias in LLMs' Economic Causal Reasoning

DGX agent

arXiv:2604.21334v1 Announce Type: new Abstract: Do large language models (LLMs) exhibit systematic ideological bias when reasoning about economic causal effects? As LLMs are increasingly used in polic

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Language as a Latent Variable for Reasoning Optimization

DGX agent

arXiv:2604.21593v1 Announce Type: new Abstract: As LLMs reduce English-centric bias, a surprising trend emerges: non-English responses sometimes outperform English on reasoning tasks. We hypothesize t

model-releasesarxiv-cs-cl
24 Apr 2026
Model Releases

Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery

DGX agent

arXiv:2604.21102v1 Announce Type: cross Abstract: We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLM

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations

DGX agent

arXiv:2509.25868v3 Announce Type: replace Abstract: The mechanisms underlying scientific confabulation in Large Language Models (LLMs) remain poorly understood. We introduce ReFACT (Reddit False And C

model-releasesarxiv-cs-cl
24 Apr 2026
Research

Slot Machines: How LLMs Keep Track of Multiple Entities

DGX agent

arXiv:2604.21139v1 Announce Type: new Abstract: Language models must bind entities to the attributes they possess and maintain several such binding relationships within a context. We study how multipl

researcharxiv-cs-cl
24 Apr 2026
Model Releases

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers

DGX agent

arXiv:2604.21700v1 Announce Type: cross Abstract: The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studie

model-releasesarxiv-cs-ai
24 Apr 2026
Tutorials

Strategic Scaling of Test-Time Compute: A Bandit Learning Approach

DGX agent

arXiv:2506.12721v2 Announce Type: replace Abstract: Scaling test-time compute has emerged as an effective strategy for improving the performance of large language models. However, existing methods typ

tutorialsarxiv-cs-ai
24 Apr 2026
Model Releases

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning

DGX agent

arXiv:2604.21327v1 Announce Type: cross Abstract: Test-time reinforcement learning (TTRL) always adapts models at inference time via pseudo-labeling, leaving it vulnerable to spurious optimization sig

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs

DGX agent

arXiv:2604.21911v1 Announce Type: cross Abstract: Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders

DGX agent

arXiv:2604.19974v1 Announce Type: cross Abstract: Large language models can be uncertain yet correct, or confident yet wrong, raising the question of whether their output-level uncertainty and their a

model-releasesarxiv-cs-cl
23 Apr 2026
Safety

Best Policy Learning from Trajectory Preference Feedback

DGX agent

arXiv:2501.18873v4 Announce Type: replace Abstract: Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful approach for aligning generative models, but its reliance on learned rew

safetyarxiv-cs-lg
23 Apr 2026
Research

Beyond ZOH: Advanced Discretization Strategies for Vision Mamba

DGX agent

arXiv:2604.20606v1 Announce Type: cross Abstract: Vision Mamba, as a state space model (SSM), employs a zero-order hold (ZOH) discretization, which assumes that input signals remain constant between s

researcharxiv-cs-ai
23 Apr 2026
Model Releases

CRAFT: Training-Free Cascaded Retrieval for Tabular QA

DGX agent

arXiv:2505.14984v2 Announce Type: replace Abstract: Open-Domain Table Question Answering (TQA) involves retrieving relevant tables from a large corpus to answer natural language queries. Traditional d

model-releasesarxiv-cs-cl
23 Apr 2026
Agents

DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data

DGX agent

arXiv:2604.19859v1 Announce Type: cross Abstract: Edge-scale deep research agents based on small language models are attractive for real-world deployment due to their advantages in cost, latency, and

agentsarxiv-cs-ai
23 Apr 2026
Model Releases

Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing

DGX agent

arXiv:2604.20429v1 Announce Type: new Abstract: Remote sensing (RS) image-text retrieval plays a critical role in understanding massive RS imagery. However, the dense multi-object distribution and com

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

Knowledge Capsules: Structured Nonparametric Memory Units for LLMs

DGX agent

arXiv:2604.20487v1 Announce Type: cross Abstract: Large language models (LLMs) encode knowledge in parametric weights, making it costly to update or extend without retraining. Retrieval-augmented gene

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Mitigating Prompt-Induced Cognitive Biases in General-Purpose AI for Software Engineering

DGX agent

arXiv:2604.16756v2 Announce Type: replace-cross Abstract: Prompt-induced cognitive biases are changes in a general-purpose AI (GPAI) system's decisions caused solely by biased wording in the input (e.

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL

DGX agent

arXiv:2604.20835v1 Announce Type: new Abstract: Modern language models demonstrate impressive coding capabilities in common programming languages (PLs), such as C++ and Python, but their performance i

model-releasesarxiv-cs-cl
23 Apr 2026
Research

R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs

DGX agent

arXiv:2604.20696v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have demonstrated impressive performance in various multimodal understanding and reasoning tasks. However, they sti

researcharxiv-cs-cv
23 Apr 2026
Tutorials

Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback

DGX agent

arXiv:2604.20730v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. Howev

tutorialsarxiv-cs-cv
23 Apr 2026
Research

Semantic-Fast-SAM: Efficient Semantic Segmenter

DGX agent

arXiv:2604.20169v1 Announce Type: new Abstract: We propose Semantic-Fast-SAM (SFS), a semantic segmentation framework that combines the Fast Segment Anything model with a semantic labeling pipeline to

researcharxiv-cs-cv
23 Apr 2026
Model Releases

SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation

DGX agent

arXiv:2604.20842v1 Announce Type: cross Abstract: Paralinguistic cues are essential for natural human-computer interaction, yet their evaluation in Large Audio-Language Models (LALMs) remains limited

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

SweRank: Software Issue Localization with Code Ranking

DGX agent

arXiv:2505.07849v2 Announce Type: replace-cross Abstract: Software issue localization, the task of identifying the precise code locations (files, classes, or functions) relevant to a natural language

model-releasesarxiv-cs-ai
23 Apr 2026
Tutorials

Transformers Can Learn Connectivity in Some Graphs but Not Others

DGX agent

arXiv:2509.22343v2 Announce Type: replace-cross Abstract: Reasoning capability is essential to ensure the factual correctness of the responses of transformer-based Large Language Models (LLMs), and ro

tutorialsarxiv-cs-ai
23 Apr 2026
Model Releases

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison

DGX agent

arXiv:2603.13779v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive success in natural visual understanding, yet they consistently underperform

model-releasesarxiv-cs-ai
22 Apr 2026
Safety

ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

DGX agent

arXiv:2604.18789v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an im

safetyarxiv-cs-ai
22 Apr 2026
Research

BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design

DGX agent

arXiv:2508.21184v3 Announce Type: replace-cross Abstract: We propose a general-purpose approach for improving the ability of large language models (LLMs) to intelligently and adaptively gather informa

researcharxiv-cs-ai
22 Apr 2026
Safety

Benchmarking Misuse Mitigation Against Covert Adversaries

DGX agent

arXiv:2506.06414v2 Announce Type: replace-cross Abstract: Existing language model safety evaluations focus on overt attacks and low-stakes tasks. In reality, an attacker can easily subvert existing sa

safetyarxiv-cs-ai
22 Apr 2026
Model Releases

Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks

DGX agent

arXiv:2512.22673v3 Announce Type: replace Abstract: Travel planning is a natural real-world task to test large language models' (LLMs) planning and tool-use abilities. Although prior work has studied

model-releasesarxiv-cs-ai
22 Apr 2026
Safety

Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs

DGX agent

arXiv:2601.15755v3 Announce Type: replace Abstract: Large language models are increasingly used to represent human opinions, values, or beliefs, and their steerability towards these ideals is an activ

safetyarxiv-cs-cl
22 Apr 2026
Model Releases

Characterizing AlphaEarth Embedding Geometry for Agentic Environmental Reasoning

DGX agent

arXiv:2604.18715v1 Announce Type: cross Abstract: Earth observation foundation models encode land surface information into dense embedding vectors, yet the geometric structure of these representations

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Deep Supervised Contrastive Learning of Pitch Contours for Robust Pitch Accent Classification in Seoul Korean

DGX agent

arXiv:2604.19477v1 Announce Type: cross Abstract: The intonational structure of Seoul Korean has been defined with discrete tonal categories within the Autosegmental-Metrical model of intonational pho

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning

DGX agent

arXiv:2604.19459v1 Announce Type: new Abstract: Formal verification guarantees proof validity but not formalization faithfulness. For natural-language logical reasoning, where models construct axiom s

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Environmental Sound Deepfake Detection Using Deep-Learning Framework

DGX agent

arXiv:2604.19652v1 Announce Type: cross Abstract: In this paper, we propose a deep-learning framework for environmental sound deepfake detection (ESDD) -- the task of identifying whether the sound sce

model-releasesarxiv-cs-ai
22 Apr 2026
Local Ai

Evaluation-driven Scaling for Scientific Discovery

DGX agent

arXiv:2604.19341v1 Announce Type: cross Abstract: Language models are increasingly used in scientific discovery to generate hypotheses, propose candidate solutions, implement systems, and iteratively

local-aiarxiv-cs-ai
22 Apr 2026
← Previous
1…329330331332333…1065
Next →