AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,376
  • Agents7,554
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,170
  • Local Ai4,930
  • Model Releases23,883
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,376
  • Agents7,554
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,170
  • Local Ai4,930
  • Model Releases23,883
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

88,376Total entries
1Added by human
88,375Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,602 results
12 May 2026

Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training

TutorialsDGX agent

arXiv:2603.04415v2 Announce Type: replace Abstract: Reasoning post-training improves Large Language Models (LLMs) on complex tasks such as mathematics and coding, but its benefits across diverse multi

Dynamic Linear Coregionalization for Realistic Synthetic Multivariate Time Series

ResearchDGX agent

arXiv:2604.05064v2 Announce Type: replace-cross Abstract: Synthetic data is essential for training foundation models for time series (FMTS), but most generators assume static correlations, and are typ

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies

Model ReleasesDGX agent

arXiv:2602.09514v3 Announce Type: replace-cross Abstract: Long-horizon planning is widely recognized as a core capability of autonomous LLM-based agents; however, current evaluation frameworks suffer

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation

Model ReleasesDGX agent

arXiv:2605.09378v1 Announce Type: cross Abstract: Long-horizon video generation has advanced in visual quality, yet existing methods still struggle to maintain knowledge consistency and coherent pedag

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

Model ReleasesDGX agent

arXiv:2605.09874v1 Announce Type: cross Abstract: Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.09826v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to track others epistemic state, makes humans efficient collaborators. AI agents need the same capacity in multi agent

Equivariant Volumetric Grasping

ApplicationsDGX agent

arXiv:2507.18847v3 Announce Type: replace-cross Abstract: We propose a new volumetric grasp model that is equivariant to rotations around the vertical axis, leading to a significant improvement in sam

Evolving Knowledge Distillation for Lightweight Neural Machine Translation

ResearchDGX agent

arXiv:2605.09924v1 Announce Type: new Abstract: Recent advancements in Neural Machine Translation (NMT) have significantly improved translation quality. However, the increasing size and complexity of

Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents

ResearchDGX agent

arXiv:2605.10663v1 Announce Type: new Abstract: Experience-driven self-evolving agents aim to overcome the static nature of large language models by distilling reusable experience from past interactio

FinTSB: A Comprehensive and Practical Benchmark for Financial Time Series Forecasting

Model ReleasesDGX agent

arXiv:2502.18834v2 Announce Type: replace-cross Abstract: Financial time series (FinTS) record the behavior of human-brain-augmented decision-making, capturing valuable historical information that can

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation

SafetyDGX agent

arXiv:2605.09430v1 Announce Type: new Abstract: Large-scale autoregressive models have demonstrated remarkable capabilities in image generation. However, their sequential raster-scan decoding relies o

From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages

Model ReleasesDGX agent

arXiv:2605.09147v1 Announce Type: cross Abstract: Part-of-speech (POS) tagging for Medieval Romance languages remains challenging due to orthographic variation, morphological complexity, and limited a

GraphBench: Next-generation graph learning benchmarking

Model ReleasesDGX agent

arXiv:2512.04475v5 Announce Type: replace-cross Abstract: Machine learning on graphs has made substantial progress across domains such as molecular property prediction and chip design. Yet benchmarkin

GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression

AgentsDGX agent

arXiv:2605.09100v1 Announce Type: new Abstract: Text embedding and generative tasks are usually trained separately based on large language models (LLMs) nowadays. This causes a large amount of trainin

GRIT: Teaching MLLMs to Think with Images

ResearchDGX agent

arXiv:2505.15879v2 Announce Type: replace-cross Abstract: Recent studies have demonstrated the efficacy of using Reinforcement Learning (RL) in building reasoning models that articulate chains of thou

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors

ResearchDGX agent

arXiv:2605.08817v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) recently thrives in large language model (LLM) reasoning tasks. However, the reward sparsity and t

Hystar: Hypernetwork-driven Style-adaptive Retrieval via Dynamic SVD Modulation

Model ReleasesDGX agent

arXiv:2605.10009v1 Announce Type: new Abstract: Query-based image retrieval (QBIR) requires retrieving relevant images given diverse and often stylistically heterogeneous queries, such as sketches, ar

Insider Attacks in Multi-Agent LLM Consensus Systems

AgentsDGX agent

arXiv:2605.08268v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in multi-agent systems where agents communicate in natural language to solve tasks jointly. A k

Investigating Anisotropy in Visual Grounding under Controlled Counterfactual Perturbations

SafetyDGX agent

arXiv:2605.09090v1 Announce Type: cross Abstract: Visual Grounding benchmarks assume that the object described by a referring expression is always present in the image, and grounding models are theref

IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts

Model ReleasesDGX agent

arXiv:2605.08664v1 Announce Type: new Abstract: Current image quality assessment methods are heavily biased towards global distortions (e.g., noise, blur), neglecting local perceptual artifacts such a

Learning Agile Striker Skills for Humanoid Soccer Robots from Noisy Sensory Input

Model ReleasesDGX agent

arXiv:2512.06571v3 Announce Type: replace Abstract: Learning fast and robust ball-kicking skills is a critical capability for humanoid soccer robots, yet it remains a challenging problem due to the ne

Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning

SafetyDGX agent

arXiv:2602.17546v2 Announce Type: replace Abstract: Instruction-following language models are trained to be helpful and safe, yet their safety behavior can deteriorate under benign fine-tuning and wor

MaD Physics: Evaluating information seeking under constraints in physical environments

Model ReleasesDGX agent

arXiv:2605.10820v1 Announce Type: new Abstract: Scientific discovery is fundamentally a resource-constrained process that requires navigating complex trade-offs between the quality and quantity of mea

MESD: A Risk-Sensitive Metric for Explanation Fairness Across Intersectional Subgroups

Model ReleasesDGX agent

arXiv:2603.13452v2 Announce Type: replace Abstract: Fairness in machine learning is predominantly evaluated through outcome-oriented metrics, such as Demographic parity, which measure whether predicti

MOTOR-Bench: A Real-world Dataset and Multi-agent Framework for Zero-shot Human Mental State Understanding

Model ReleasesDGX agent

arXiv:2605.09703v1 Announce Type: new Abstract: Understanding human mental states from natural behavior is crucial for intelligent systems in the real world. However, most current research focuses on

Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy

Model ReleasesDGX agent

arXiv:2605.10550v1 Announce Type: new Abstract: Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -

Multi-layer attentive probing improves transfer of audio representations for bioacoustics

SafetyDGX agent

arXiv:2605.10494v1 Announce Type: cross Abstract: Probing heads map the representations learned from audio by a machine learning model to downstream task labels and are a key component in evaluating r

Narrative Landscape: Mapping Narrative Dispositions Across LLMs

ResearchDGX agent

arXiv:2605.08742v1 Announce Type: cross Abstract: This study proposes a quantitative framework for profiling LLM dispositions as stable, model-specific regularities in output under repeated, controlle

Neuroprobe: Evaluating Intracranial Brain Responses to Naturalistic Stimuli

ResearchDGX agent

arXiv:2509.21671v2 Announce Type: replace Abstract: High-resolution neural datasets enable foundation models for the next generation of brain-computer interfaces and neurological treatments. The commu

NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces

ApplicationsDGX agent

arXiv:2510.03895v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world de

On Variance Reduction in Learning Mean Flows

SafetyDGX agent

arXiv:2605.09235v1 Announce Type: cross Abstract: One-step generative modeling has emerged as a leading approach to amortize the inference cost of diffusion and flow-matching models. Among distillatio

OracleTSC: Oracle-Informed Reward Hurdle and Uncertainty Regularization for Traffic Signal Control

Model ReleasesDGX agent

arXiv:2605.08516v1 Announce Type: new Abstract: Transparent decision-making is essential for traffic signal control (TSC) systems to earn public trust. However, traditional reinforcement learning-base

ORICF -- Open Robotics Inference and Control Framework

ResearchDGX agent

arXiv:2605.09656v1 Announce Type: new Abstract: Recent advances in artificial intelligence (AI) have enabled effective perception and language models for robots, but their deployment remains computati

Parallel Multi-Circuit Quantum Feature Fusion in Hybrid Quantum-Classical Convolutional Neural Networks for Breast Tumor Classification

Model ReleasesDGX agent

arXiv:2512.02066v2 Announce Type: replace-cross Abstract: Quantum machine learning has emerged as a promising approach to improve feature extraction and classification tasks in high-dimensional data d

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs

SafetyDGX agent

arXiv:2605.09422v1 Announce Type: new Abstract: Although Large Multimodal Models (LMMs) have achieved strong performance on general video understanding, their susceptibility to textual prior shortcuts

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

SafetyDGX agent

arXiv:2604.01532v2 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on

PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting

AgentsDGX agent

arXiv:2605.08935v1 Announce Type: new Abstract: Coupled spatiotemporal forecasting is important for predicting the future evolution of multiple interacting dynamical systems, such as in climate models

Predictive Radiomics for Evaluation of Cancer Immune SignaturE in Glioblastoma: the PRECISE-GBM study

ResearchDGX agent

arXiv:2605.10278v1 Announce Type: new Abstract: Background: Radiogenomics allows identification of radiological biomarkers for genomic phenotypes. In glioblastoma, these biomarkers could potentially c

Priority-Driven Control and Communication in Decentralized Multi-Agent Systems via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.10482v1 Announce Type: cross Abstract: Event-triggered control provides a mechanism for avoiding excessive use of constrained communication bandwidth in networked multi-agent systems. Howev

ReaMOT: A Benchmark and Framework for Reasoning-based Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2505.20381v4 Announce Type: replace Abstract: Referring Multi-Object Tracking (RMOT) aims to track targets specified by language instructions. However, existing RMOT paradigms heavily rely on ex

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

Model ReleasesDGX agent

arXiv:2604.01527v3 Announce Type: replace-cross Abstract: Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fi

REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?

Model ReleasesDGX agent

arXiv:2505.10872v4 Announce Type: replace-cross Abstract: Robot task planning decomposes human instructions into executable action sequences that enable robots to complete a series of complex tasks. A

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

Model ReleasesDGX agent

arXiv:2605.10921v1 Announce Type: new Abstract: Memory is a critical component of robotic intelligence, as robots must rely on past observations and actions to accomplish long-horizon tasks in partial

RuPLaR : Efficient Latent Compression of LLM Reasoning Chains with Rule-Based Priors From Multi-Step to One-Step

SafetyDGX agent

arXiv:2605.09346v1 Announce Type: cross Abstract: The Chain-of-Thought (CoT) paradigm, while enhancing the interpretability of Large Language Models (LLMs), is constrained by the inefficiencies and ex

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems

Model ReleasesDGX agent

arXiv:2605.10246v1 Announce Type: new Abstract: AI scientist systems are increasingly deployed for autonomous research, yet their academic integrity has never been systematically evaluated. We introdu

SciLT: Long-tailed Image Classification under Scientific Image Domains

ResearchDGX agent

arXiv:2604.03687v2 Announce Type: replace Abstract: Long-tailed recognition has benefited from foundation models and fine-tuning paradigms, yet existing studies and benchmarks are mainly confined to n

Semi-Supervised Neural Super-Resolution for Mesh-Based Simulations

Model ReleasesDGX agent

arXiv:2605.09284v1 Announce Type: cross Abstract: Mesh-based simulations provide high-fidelity solutions to partial differential equations (PDEs), but achieving such accuracy typically requires fine m

Set Prediction for Next-Day Active Fire Forecasting

Model ReleasesDGX agent

arXiv:2605.10298v1 Announce Type: new Abstract: Accurate next-day active fire forecasts can support early warning, disaster response, forest risk assessment, and downstream estimation of fire-related

simpleposter: a simple baseline for product poster generation

Model ReleasesDGX agent

arXiv:2605.08784v1 Announce Type: new Abstract: Product poster generation poses distinct challenges beyond general poster design, requiring both faithful preservation of product appearance and precise

SpectraLLM: Uncovering the Ability of LLMs for Molecular Structure Elucidation from Multi-Spectral Data

Model ReleasesDGX agent

arXiv:2508.08441v3 Announce Type: replace-cross Abstract: Automated molecular structure elucidation remains challenging, as existing approaches often depend on pre-compiled databases or restrict thems

Step 3.5 Flash from @StepFun_ai is now free again in Nous Portal for the next 15 days!

ResearchDGX agent

Nous Research announced that Step 3.5 Flash, an AI model from StepFun, is temporarily available for free on the Nous Portal for a 15-day period. This offer provides users access to the model without c

Structured Recurrent Mixers for Massively Parallelized Sequence Generation

ResearchDGX agent

arXiv:2605.08696v1 Announce Type: new Abstract: Over the last two decades, language modeling has experienced a shift from predominantly recurrent architectures that process tokens sequentially during

Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions

ResearchDGX agent

arXiv:2605.09967v1 Announce Type: new Abstract: While researchers are finding concepts represented as linear directions in language models, a bag of linear directions fails to capture relational struc

TextBridgeGNN: Pre-training Graph Neural Network for Cross-Domain Recommendation via Text-Guided Transfer

TutorialsDGX agent

arXiv:2601.02366v2 Announce Type: replace-cross Abstract: Graph-based recommendation has achieved great success in recent years. The classical graph recommendation model utilizes ID embedding to store

The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection

Model ReleasesDGX agent

arXiv:2605.10334v1 Announce Type: new Abstract: Recent deepfake detection methods demonstrate improved cross-dataset generalization, yet the underlying mechanisms remain underexplored. We introduce th

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space

ResearchDGX agent

arXiv:2605.09883v1 Announce Type: cross Abstract: As current Multimodal Large Language Models rapidly saturate canonical visual reasoning benchmarks, a key question emerges: do these strong scores gen

The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans

SafetyDGX agent

arXiv:2605.08837v1 Announce Type: cross Abstract: Abstract concepts - justice, theory, availability - have no single perceivable referent; in the human brain, their meaning emerges from a web of exper

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime

ResearchDGX agent

arXiv:2605.10601v1 Announce Type: new Abstract: AI deployment in sensitive domains such as health care, credit, employment, and criminal justice is often treated as unsafe to authorize until model int

ThreatCore: A Benchmark for Explicit and Implicit Threat Detection

Model ReleasesDGX agent

arXiv:2605.10563v1 Announce Type: cross Abstract: Threat detection in Natural Language Processing lacks consistent definitions and standardized benchmarks, and is often conflated with broader phenomen

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

SafetyDGX agent

arXiv:2508.20697v3 Announce Type: replace-cross Abstract: As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studie

← Previous
1…476477478479480…1061
Next →