AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
Human
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
59,851 results
29 Apr 2026

Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories

ApplicationsDGX agent

arXiv:2411.05174v2 Announce Type: replace Abstract: We consider the problem of estimating the transition dynamics T^* from near-optimal expert trajectories in the context of offline model-based reinfo

Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance

Model ReleasesDGX agent

arXiv:2604.25249v1 Announce Type: new Abstract: Detecting sandbagging--the deliberate underperformance on capability evaluations--is an open problem in AI safety. We tested whether symptom validity te

BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

Model ReleasesDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.24955v1 Announce Type: new Abstract: As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken

Benchmarking and Adapting On-Device LLMs for Clinical Decision Support

Model ReleasesDGX agent

arXiv:2601.03266v2 Announce Type: replace Abstract: Large language models (LLMs) have rapidly advanced in clinical decision-making, yet the deployment of proprietary systems is hindered by privacy con

Benchmarking and Improving GUI Agents in High-Dynamic Environments

Model ReleasesDGX agent

arXiv:2604.25380v1 Announce Type: new Abstract: Recent advancements in Graphical User Interface (GUI) agents have predominantly focused on training paradigms like supervised fine-tuning (SFT) and rein

Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings

Model ReleasesDGX agent

arXiv:2604.25358v1 Announce Type: new Abstract: Evaluating layout-guided text-to-image generative models requires assessing both semantic alignment with textual prompts and spatial fidelity to prescri

Benchmarking Logistic Regression, SVM, and LightGBM Against BiLSTM with Attention for Sentiment Analysis on Indonesian Product Reviews

ResearchDGX agent

arXiv:2604.25452v1 Announce Type: new Abstract: Sentiment analysis of product reviews on e-commerce platforms plays a critical role in automatically understanding customer satisfaction and providing a

Benchmarking OCR Pipelines with Adaptive Enhancement for Multi-Domain Retail Bill Digitization

Model ReleasesDGX agent

arXiv:2604.25176v1 Announce Type: new Abstract: The digitization of multi-domain retail billing documents remains a challenging task due to variability in scan quality, layout heterogeneity, and domai

Benchmarking PyCaret AutoML Against IndoBERT Fine-Tuning for Sentiment Analysis on Indonesian IKN Twitter Data

ResearchDGX agent

arXiv:2604.25392v1 Announce Type: new Abstract: This paper benchmarks a classical machine learning approach based on PyCaret AutoML against a deep learning approach based on IndoBERT fine-tuning for b

BEVal: A Cross-dataset Evaluation Study of BEV Segmentation Models for Autonomous Driving

AgentsDGX agent

arXiv:2408.16322v4 Announce Type: replace Abstract: Current research in semantic bird's-eye view segmentation for autonomous driving focuses solely on optimizing neural network models using a single d

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

SafetyDGX agent

arXiv:2604.25072v1 Announce Type: new Abstract: Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evalua

Beyond Fidelity: Semantic Similarity Assessment in Low-Level Image Processing

ResearchDGX agent

arXiv:2604.25408v1 Announce Type: new Abstract: Low-level image processing has long been evaluated mainly from the perspective of visual fidelity. However, with the rise of deep learning and generativ

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal

Model ReleasesDGX agent

arXiv:2509.09708v3 Announce Type: replace Abstract: Refusal on harmful prompts is a key safety behaviour in instruction-tuned large language models (LLMs), yet the internal causes of this behaviour re

Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Space Models

ResearchDGX agent

arXiv:2604.25416v1 Announce Type: new Abstract: Model-Based Reinforcement Learning distinguishes between physical dynamics models operating on proprioceptive inputs and latent dynamics models operatin

BifDet: A 3D Bifurcation Detection Dataset for Airway-Tree Modeling

Model ReleasesDGX agent

arXiv:2604.24999v1 Announce Type: new Abstract: Thoracic Computed Tomography (CT) scans offer detailed insights into the intricate branching network of the airway tree, which is essential for understa

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding

Model ReleasesDGX agent

arXiv:2512.12087v3 Announce Type: replace Abstract: The growing demand for long-context inference capabilities in Large Language Models (LLMs) has intensified the computational and memory bottlenecks

Bridging the Indoor-Outdoor Gap: Cross-Technology Ranging for Seamless Robot Navigation

ResearchDGX agent

arXiv:2604.25541v1 Announce Type: cross Abstract: Mobile robots that move between outdoor and indoor environments still struggle with consistent positioning. Satellite-based and terrestrial ranging ea

Bug-Report-Driven Fault Localization: Industrial Benchmarking and Lesson Learned at ABB Robotics

Local AiDGX agent

arXiv:2604.25700v1 Announce Type: cross Abstract: Software quality assurance remains a major challenge in industrial environments, where large-scale and long-lived systems inevitably accumulate defect

Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation

ResearchDGX agent

arXiv:2604.25580v1 Announce Type: new Abstract: The closure of Perspective API at the end of 2026 discards what has functioned as the de facto standard for automated toxicity measurement in NLP, CSS,

C3G: Learning Compact 3D Representations with 2K Gaussians

TutorialsDGX agent

arXiv:2512.04021v2 Announce Type: replace Abstract: Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. R

Calibrated Fusion for Heterogeneous Graph-Vector Retrieval in Multi-Hop QA

ResearchDGX agent

arXiv:2603.28886v2 Announce Type: replace-cross Abstract: Graph-augmented retrieval combines dense similarity with graph-based relevance signals such as Personalized PageRank (PPR), but these scores h

CAN-QA: A Question-Answering Benchmark for Reasoning over In-Vehicle CAN Traffic

Model ReleasesDGX agent

arXiv:2604.24935v1 Announce Type: cross Abstract: The Controller Area Network (CAN) is a safety-critical in-vehicle communication protocol that lacks built-in security mechanisms, making intrusion det

Can We Change the Stroke Size for Easier Diffusion?

ResearchDGX agent

arXiv:2603.26783v2 Announce Type: replace Abstract: Diffusion models can be challenged in the low signal-to-noise regime, where they have to make pixel-level predictions despite the presence of high n

Carbon-Taxed Transformers: A Green Compression Pipeline for Overgrown Language Models

SafetyDGX agent

arXiv:2604.25903v1 Announce Type: cross Abstract: The accelerating adoption of Large Language Models (LLMs) in software engineering (SE) has brought with it a silent crisis: unsustainable computationa

Categorical Optimization with Bayesian Anchored Latent Trust Regions for Structural Design under High-Dimensional Uncertainty

ResearchDGX agent

arXiv:2604.25241v1 Announce Type: new Abstract: Categorical structural optimization under aleatoric uncertainty is challenging because each design variable must be selected from a finite catalog of ad

CGU-ILALab at FoodBench-QA 2026: Comparing Traditional and LLM-based Approaches for Recipe Nutrient Estimation

Model ReleasesDGX agent

arXiv:2604.25774v1 Announce Type: new Abstract: Accurate nutrient estimation from unstructured recipe text is an important yet challenging problem in dietary monitoring, due to ambiguous ingredient te

Cheaper, Better, Faster, Stronger: Robust Text-to-SQL without Chain-of-Thought or Fine-Tuning

Model ReleasesDGX agent

arXiv:2505.14174v2 Announce Type: replace Abstract: LLMs are effective at code generation tasks like text-to-SQL, but is it worth the cost? Many state-of-the-art approaches use non-task-specific LLM t

CHUCKLE -- When Humans Teach AI To Learn Emotions The Easy Way

SafetyDGX agent

arXiv:2510.09382v2 Announce Type: replace Abstract: Curriculum learning (CL) structures training from simple to complex samples, facilitating progressive learning. However, existing CL approaches for

Citation Failure: Definition, Analysis and Efficient Mitigation

Model ReleasesDGX agent

arXiv:2510.20303v3 Announce Type: replace Abstract: Citations from LLM-based RAG systems are supposed to simplify response verification. However, this goal is undermined in cases of citation failure,

CiteRadar: A Citation Intelligence Platform for Researcher Profiling and Geographic Visualization

ResearchDGX agent

arXiv:2604.25057v1 Announce Type: new Abstract: Understanding the geographic reach and community structure of one's scholarly citations is increasingly valuable for career development, grant applicati

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding

ResearchDGX agent

arXiv:2602.01785v2 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable success in source code understanding, yet as software systems grow in scale, computational eff

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval

Model ReleasesDGX agent

arXiv:2604.25273v1 Announce Type: new Abstract: Despite significant progress in Unified Multimodal Retrieval (UMR) powered by Large Multimodal Models (LMMs), existing embedding methods primarily focus

Comparative Study of Bending Analysis using Physics-Informed Neural Networks and Numerical Dynamic Deflection in Perforated nanobeam

ResearchDGX agent

arXiv:2604.24768v1 Announce Type: new Abstract: In this chapter, we investigate the bending behavior of a perforated nanobeam subjected to sinusoidal loading using an efficient and computationally rob

Comparing Data Assimilation and Likelihood-Based Inference on Latent State Estimation in Agent-Based Models

Model ReleasesDGX agent

arXiv:2509.17625v2 Announce Type: replace Abstract: In this paper, we present the first systematic comparison of Data Assimilation (DA) and Likelihood-Based Inference (LBI) in the context of an Agent-

COMPASS: COmpact Multi-channel Prior-map And Scene Signature for Floor-Plan-Based Visual Localization

Local AiDGX agent

arXiv:2604.25388v1 Announce Type: new Abstract: Architectural floor plans are widely available priors which contain not only geometry but also the semantic information of the environment, yet existing

Compute Aligned Training: Optimizing for Test Time Inference

SafetyDGX agent

arXiv:2604.24957v1 Announce Type: new Abstract: Scaling test-time compute has emerged as a powerful mechanism for enhancing Large Language Model (LLM) performance. However, standard post-training para

Conditional Flow Matching for Probabilistic Downscaling of Maximum 3-day Snowfall in Alaska

ResearchDGX agent

arXiv:2604.25172v1 Announce Type: cross Abstract: Precipitation in complex terrain is governed by orographic processes operating at scales of a few kilometers, yet climate models typically run at reso

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

SafetyDGX agent

arXiv:2604.25891v1 Announce Type: new Abstract: Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavio

Contrast-Enhanced Gating in GRUs for Robust Low-Data Sequence Learning

Model ReleasesDGX agent

arXiv:2402.09034v3 Announce Type: replace Abstract: Activation functions govern how recurrent networks regulate and transmit information across temporal dependencies. Despite advances in sequence mode

Contrastive Image-Metadata Pre-Training for Materials Transmission Electron Microscopy

TutorialsDGX agent

arXiv:2604.24909v1 Announce Type: new Abstract: The vast majority of transmission electron microscopy (TEM) data never gets published and ends up on a backup drive until deleted to free up space. Thes

Control Your Queries: Heterogeneous Query Interaction for Camera-Radar Fusion

AgentsDGX agent

arXiv:2604.25574v1 Announce Type: new Abstract: In autonomous driving, camera-radar fusion offers complementary sensing and low deployment cost. Existing methods perform fusion through input mixing, f

Cooperate to Compete: Strategic Coordination in Multi-Agent Conquest

AgentsDGX agent

arXiv:2604.25088v1 Announce Type: cross Abstract: Language Model (LM)-based agents remain largely untested in mixed-motive settings where agents must leverage short-term cooperation for long-term comp

CORAL: Adaptive Retrieval Loop for Culturally-Aligned Multilingual RAG

SafetyDGX agent

arXiv:2604.25676v1 Announce Type: new Abstract: Multilingual retrieval-augmented generation (mRAG) is often implemented within a fixed retrieval space, typically via query or document translation or m

CoRE: Concept-Reasoning Expansion for Continual Brain Lesion Segmentation

Model ReleasesDGX agent

arXiv:2604.25376v1 Announce Type: new Abstract: Accurate brain lesion segmentation in MRI is vital for effective clinical diagnosis and treatment planning. Due to high annotation costs and strict data

CoreFlow: Low-Rank Matrix Generative Models

ResearchDGX agent

arXiv:2604.24959v1 Announce Type: new Abstract: Learning matrix-valued distributions from high-dimensional and possibly incomplete training data is challenging: ambient-space generative modeling is co

Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models

ResearchDGX agent

arXiv:2603.12118v2 Announce Type: replace Abstract: Any-to-Any models are an emerging class of multimodal models that accept combinations of multimodal data (e.g., text, image, video, audio) as input

CRAFT: Grounded Multi-Agent Coordination Under Partial Information

Model ReleasesDGX agent

arXiv:2603.25268v2 Announce Type: replace Abstract: We introduce CRAFT, a multi-agent benchmark for evaluating pragmatic communication in large language models under strict partial information. In thi

CRC-SAM: SAM-Based Multi-Modal Segmentation and Quantification of Colorectal Cancer in CT, Colonoscopy, and Histology Images

ResearchDGX agent

arXiv:2604.24793v1 Announce Type: cross Abstract: We present CRC-SAM, a unified framework for colorectal cancer segmentation across colonoscopy, CT, and histopathology images. Unlike prior single-moda

CroSearch-R1: Better Leveraging Cross-lingual Knowledge for Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2604.25182v1 Announce Type: new Abstract: A multilingual collection may contain useful knowledge in other languages to supplement and correct the facts in the original language for Retrieval-Aug

Cross-Lingual Jailbreak Detection via Semantic Codebooks

Model ReleasesDGX agent

arXiv:2604.25716v1 Announce Type: new Abstract: Safety mechanisms for large language models (LLMs) remain predominantly English-centric, creating systematic vulnerabilities in multilingual deployment.

Curl Descent: Non-Gradient Learning Dynamics with Sign-Diverse Plasticity

ResearchDGX agent

arXiv:2510.02765v4 Announce Type: replace Abstract: Gradient-based algorithms are a cornerstone of artificial neural network training, yet it remains unclear whether biological neural networks use sim

Curriculum-guided multimodal representation learning enables generalizable prediction of nanomaterial-protein interactions

ResearchDGX agent

arXiv:2507.14245v2 Announce Type: replace Abstract: Nanomaterial-protein interactions (NPI) are pivotal to realizing the therapeutic and diagnostic potential of nanomaterials. Although AI promises to

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation

Model ReleasesDGX agent

arXiv:2604.25318v1 Announce Type: cross Abstract: Cutscenes are carefully choreographed cinematic sequences embedded in video games and interactive media, serving as the primary vehicle for narrative

Data-Driven Hamiltonian Reduction for Superconducting Qubits via Meta-Learning

Model ReleasesDGX agent

arXiv:2604.24912v1 Announce Type: cross Abstract: We introduce HAML (Hamiltonian Adaptation via Meta-Learning), a framework for fast online adaptation of effective Hamiltonian models of superconductin

DCD: Decomposition-based Causal Discovery from Autocorrelated and Non-Stationary Temporal Data

ApplicationsDGX agent

arXiv:2602.01433v2 Announce Type: replace Abstract: Multivariate time series in domains such as finance, climate science, and healthcare often exhibit long-term trends, seasonal patterns, and short-te

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

TutorialsDGX agent

arXiv:2604.25477v1 Announce Type: new Abstract: Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance t

Deflation-Free Optimal Scoring

ApplicationsDGX agent

arXiv:2604.25664v1 Announce Type: cross Abstract: Sparse Optimal Scoring (SOS) reformulates linear discriminant analysis to enable feature selection through elastic net regularization, making it well-

DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework

SafetyDGX agent

arXiv:2506.05199v3 Announce Type: replace Abstract: A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pa

DenseScout: Algorithm-System Co-design for Budgeted Tiny Object Selection on Edge Platforms

ResearchDGX agent

arXiv:2604.25300v1 Announce Type: new Abstract: Deploying tiny object perception on edge platforms is challenging because practical systems must satisfy both strict compute budgets and end-to-end late

Detecting Dental Landmarks from Intraoral 3D Scans: the 3DTeethLand challenge

Model ReleasesDGX agent

arXiv:2512.08323v2 Announce Type: replace Abstract: Teeth landmark detection is a key task in modern orthodontics, supporting advanced diagnosis, personalized treatment planning, and effective monitor

← Previous
1…815816817818819…998
Next →