AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
3 Jul 2026

Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models

Model ReleasesDGX agent

arXiv:2607.01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes va

MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction

Model ReleasesDGX agent

arXiv:2607.01627v1 Announce Type: cross Abstract: Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development. A difficul

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.01813v1 Announce Type: cross Abstract: Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to

MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support

Model ReleasesDGX agent

arXiv:2607.01814v1 Announce Type: new Abstract: Traditional Chinese Medicine (TCM) diagnosis, particularly through tongue inspection, faces persistent challenges in subjectivity and reproducibility. T

Model Merging as Probabilistic Inference in Fine-Tuning Parameter Space

Model ReleasesDGX agent

arXiv:2607.01689v1 Announce Type: cross Abstract: Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.~Most existing appr

MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding

SafetyDGX agent

arXiv:2607.01982v1 Announce Type: cross Abstract: Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a new trend in

Morphology-Aware Sample Assignment: Overcoming IoU Insensitivity for Surface Defect Detection

SafetyDGX agent

arXiv:2606.13723v2 Announce Type: replace-cross Abstract: Intersection-over-Union (IoU), as a pivotal metric for evaluating the spatial alignment between candidate proposals and ground-truth annotatio

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

Model ReleasesDGX agent

arXiv:2607.01420v1 Announce Type: cross Abstract: As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and

Multi-Head Recurrent Memory Agents

ResearchDGX agent

arXiv:2607.01523v1 Announce Type: cross Abstract: Recurrent memory agents extend LLMs to arbitrarily long contexts by iteratively consolidating input into a fixed-size memory window. Despite their sca

Multi-modal Rail Crossing Safety Analysis

SafetyDGX agent

arXiv:2607.01365v1 Announce Type: cross Abstract: Given one or more images of a railway crossing, can we leverage visual cues that allow us to robustly estimate how safe it is? Can we improve our abil

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

Model ReleasesDGX agent

arXiv:2604.04532v2 Announce Type: replace-cross Abstract: Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language

Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing

Local AiDGX agent

arXiv:2607.01978v1 Announce Type: new Abstract: Online multimodal knowledge editing requires injecting a continual stream of visual-textual corrections into multimodal large language models (MLLMs) wi

mupscaling small models: Principled warm starts and hyperparameter transfer

Model ReleasesDGX agent

arXiv:2602.10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve effic

Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging

SafetyDGX agent

arXiv:2510.17426v3 Announce Type: replace-cross Abstract: The 'alignment tax' of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration, ma

NeoMap: Training-free Novel-View Synthesis from Single Images and Videos

SafetyDGX agent

arXiv:2607.01962v1 Announce Type: cross Abstract: We study the challenging problem of novel view video synthesis from single images or monocular videos. Existing methods, which operate under the assum

NeuroBridge: Bridging Multi-Task MRI Knowledge for Neurodegenerative Disease Diagnosis

ResearchDGX agent

arXiv:2607.01401v1 Announce Type: cross Abstract: INTRODUCTION: Accurate MRI-based identification of Alzheimer's disease (AD), mild cognitive impairment (MCI), and related dementias remains challengin

Neuron-Aware Active Few-Shot Learning for LLMs

ResearchDGX agent

arXiv:2607.02423v1 Announce Type: cross Abstract: Active Few-Shot Learning (AFSL) adapts LLMs to specialized domains by identifying the most valuable unlabeled samples for annotation and use as few-sh

Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation

SafetyDGX agent

arXiv:2607.02460v1 Announce Type: cross Abstract: Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in s

Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimization

SafetyDGX agent

arXiv:2607.01972v1 Announce Type: cross Abstract: Large language models (LLMs) are often asked to produce JSON conforming to a fixed schema, powering information extraction, tool calling, agentic plan

Office Comprehension Benchmark

Model ReleasesDGX agent

arXiv:2607.01245v1 Announce Type: cross Abstract: We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension

OmniGAIA: Towards Native Omni-Modal AI Agents

Model ReleasesDGX agent

arXiv:2602.22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to i

On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain

Model ReleasesDGX agent

arXiv:2607.01444v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models offer inference speedups via selective activation but impose substantial memory requirements because the whole network

Online Safety Monitoring for LLMs

SafetyDGX agent

arXiv:2607.02510v1 Announce Type: new Abstract: Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safet

OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models

ResearchDGX agent

arXiv:2607.01977v1 Announce Type: new Abstract: Ontology learning (OL) aims to automatically construct structured knowledge models from text, yet progress remains fragmented across methods, domains, a

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

Model ReleasesDGX agent

arXiv:2607.02047v1 Announce Type: cross Abstract: Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts.

Ophiuchus: Incentivizing Tool-augmented 'Think with Images' for Joint Medical Segmentation, Understanding and Reasoning

AgentsDGX agent

arXiv:2512.14157v2 Announce Type: replace Abstract: Recent medical MLLMs have made significant progress in generating step-by-step textual reasoning chains. However, they still struggle with complex c

OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration

Model ReleasesDGX agent

arXiv:2607.01531v1 Announce Type: new Abstract: Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networ

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning

ResearchDGX agent

arXiv:2604.02091v2 Announce Type: replace-cross Abstract: Rerankers play a pivotal role in refining retrieval results for Retrieval-Augmented Generation. However, current reranking models are typicall

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

Model ReleasesDGX agent

arXiv:2607.02461v1 Announce Type: cross Abstract: Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make infe

Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond

ResearchDGX agent

arXiv:2607.02197v1 Announce Type: cross Abstract: The society and emerging risk-based regulatory frameworks for AI underscore the need for rigorous risk assessment to ensure safe and reliable AI syste

PACE: A Neuro-Symbolic Framework for Plausible and Actionable Counterfactual Explanations

ApplicationsDGX agent

arXiv:2607.01306v1 Announce Type: new Abstract: Counterfactual explanations explain machine learning predictions by identifying minimal input changes that would alter a model's decision. Although many

PACE: A Proxy for Agentic Capability Evaluation

Model ReleasesDGX agent

arXiv:2607.02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation c

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2607.01754v1 Announce Type: new Abstract: On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribu

PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

Model ReleasesDGX agent

arXiv:2607.01938v1 Announce Type: cross Abstract: Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action

Playing 20 Question Game with Policy-Based Reinforcement Learning

SafetyDGX agent

arXiv:1808.07645v5 Announce Type: replace-cross Abstract: The 20 Questions (Q20) game is a well known game which encourages deductive reasoning and creativity. In the game, the answerer first thinks o

Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack

ResearchDGX agent

arXiv:2607.01702v1 Announce Type: cross Abstract: Recently, speech classification methods have gained widespread adoption in intelligent gadgets. Current study indicates that backdoor attacks provide

Population-Based Multi-Objective Training of Discriminators for Semi-Supervised GANs

ResearchDGX agent

arXiv:2607.01907v1 Announce Type: cross Abstract: Semi-supervised generative adversarial networks (SSL-GANs) can exploit large unlabeled datasets while retaining a classifier in the discriminator, but

Power Systems Agent Benchmark: Executable Evaluation of AI Agents in Electric Power Engineering

Model ReleasesDGX agent

arXiv:2606.20950v2 Announce Type: replace Abstract: Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way

PPTArena: A Benchmark for PowerPoint Editing

Model ReleasesDGX agent

arXiv:2512.03042v3 Announce Type: replace-cross Abstract: We introduce PPTArena, a benchmark for PowerPoint editing that evaluates how agents modify real slides from natural-language instructions. Unl

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge

Model ReleasesDGX agent

arXiv:2607.01829v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing a

Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander

SafetyDGX agent

arXiv:2607.01736v1 Announce Type: cross Abstract: We study how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone. Choosing the ri

Predicting Early Stages Of Alzheimer's Disease And Identifying Key Biomarkers Using Deep Artificial Neural Network And Ensemble Of Machine Learning Methodologies

ResearchDGX agent

arXiv:2607.02142v1 Announce Type: cross Abstract: Alzheimers disease (AD) is a brain disorder that develops slowly and mainly affects memory, thinking, language, and daily activities. It is one of the

PreScience: A Dataset and Benchmark for Scientific Forecasting

Model ReleasesDGX agent

arXiv:2602.20459v2 Announce Type: replace Abstract: Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark fo

ProCal: Inference-Time Proposal Calibration for Open-Vocabulary Object Detection

Local AiDGX agent

arXiv:2607.01759v1 Announce Type: cross Abstract: Open-vocabulary object detection aims to localize and classify objects beyond the fixed set of categories seen dur ing training. Recent open-vocabular

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models

Local AiDGX agent

arXiv:2607.01480v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR), along with recent selfdistillation variants such as SDPO, evaluates each rollout against a verifi

Profit-Based Counterfactual Explanations for Product Improvement: A Case Study of Manga Sales in Japan

ApplicationsDGX agent

arXiv:2607.01610v1 Announce Type: new Abstract: Counterfactual explanation (CE) is widely used to enhance the interpretability of machine learning models and support data-driven decision-making based

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Model ReleasesDGX agent

arXiv:2607.02512v1 Announce Type: cross Abstract: Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON, or ranking

Prompt Coverage Adequacy

AgentsDGX agent

arXiv:2607.02057v1 Announce Type: cross Abstract: In recent years, it has become increasingly evident that large language models (LLMs) and autonomous agents raise the level of abstraction in software

Prompt Framing Distorts Count-Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

Model ReleasesDGX agent

arXiv:2607.01240v1 Announce Type: cross Abstract: Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding i

Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate

Model ReleasesDGX agent

arXiv:2510.04391v5 Announce Type: replace Abstract: Mental imagery vividness is a stable individual trait, yet whether imagined scenarios share relational structure across human and synthetic large la

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

Model ReleasesDGX agent

arXiv:2510.04484v2 Announce Type: replace-cross Abstract: The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactio

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

SafetyDGX agent

arXiv:2607.02234v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference

QFedAgent: Quantum-Enhanced Personalized Federated Learning for Multi-Agent Activity Recognition

Model ReleasesDGX agent

arXiv:2607.02426v1 Announce Type: cross Abstract: Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data, making it suitable for privacy-sensi

RadiomicNet: A Hybrid Radiomics-Guided Lightweight Architecture for Interpretable Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.02185v1 Announce Type: cross Abstract: Deep learning has achieved remarkable performance in medical image segmentation, yet it suffers from critical limitations: mathematical intractability

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

SafetyDGX agent

arXiv:2607.01897v1 Announce Type: cross Abstract: We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards. RTA trains a

Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study

Local AiDGX agent

arXiv:2607.02436v1 Announce Type: cross Abstract: Agentic coding assistants are increasingly given extra capabilities, such as browser based testing tools and design oriented system prompts, on the as

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

Model ReleasesDGX agent

arXiv:2607.02504v1 Announce Type: cross Abstract: Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on extbf{sp

ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning

ResearchDGX agent

arXiv:2607.02509v1 Announce Type: new Abstract: Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Althou

RedCoder: Automated Multi-Turn Red Teaming for Code LLMs

SafetyDGX agent

arXiv:2507.22063v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) for code generation (i.e., Code LLMs) have demonstrated impressive capabilities in AI-assisted software developme

Reformalization of the Jordan Curve Theorem

ApplicationsDGX agent

arXiv:2607.01734v1 Announce Type: new Abstract: We present a case study in reformalization, a variant of autoformalization in which the input proof is not natural language but a formal development in

← Previous
1…96979899100…358
Next →