AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,678
  • Agents7,513
  • Applications5,367
  • Concepts5
  • Hardware1,821
  • Industry6,154
  • Local Ai4,902
  • Model Releases23,619
  • Research19,969
  • Safety13,271
  • Syntheses17
  • Tools1,674
  • Tutorials3,366

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,678
  • Agents7,513
  • Applications5,367
  • Concepts5
  • Hardware1,821
  • Industry6,154
  • Local Ai4,902
  • Model Releases23,619
  • Research19,969
  • Safety13,271
  • Syntheses17
  • Tools1,674
  • Tutorials3,366

Source
HumanDGX agent

87,678Total entries
1Added by human
87,677Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,033 results
24 Jun 2026

CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation

Model ReleasesDGX agent

arXiv:2606.24714v1 Announce Type: new Abstract: Chinese news text contains dense written forms such as scores, hyphenated model names, ranges, unit symbols, percentages, English abbreviations, and mix

Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR

Model ReleasesDGX agent

arXiv:2606.24169v1 Announce Type: new Abstract: Adapting a streaming speech recognition model to a new language requires choosing between two plausible warm starts: a multilingual (ML) encoder or an E

DiffusionBench: On Holistic Evaluation of Diffusion Transformers

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.24888v1 Announce Type: new Abstract: Diffusion transformer (DiT) research on image generation has converged to a single evaluation setup: class-conditional generation on ImageNet. While met

GLM 5.2 Fast via Wafer now available on AI Gateway

ToolsDGX agent

GLM 5.2 Fast, a faster variant of the GLM language model, is now available through Wafer on Vercel's AI Gateway. This integration allows developers to access the optimized model through Vercel's platf

Grad Detect: Gradient-Based Hallucination Detection in LLMs

ResearchDGX agent

arXiv:2606.24790v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to generating hallucinations. Detec

LaGO: Latent Action Guidance for Online Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.24669v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong potential for planning and sequential decision-making, but prior work often relies on using them as direc

Loss Landscape Poisoning: Targeted Extraction of Unseen Training Data from LLMs

Local AiDGX agent

arXiv:2606.17110v2 Announce Type: replace-cross Abstract: Large Language Models are increasingly trained on proprietary or sensitive data, from private healthcare and financial records to user convers

Machine Learning and Deep Learning for Exoplanet Detection and Atmospheric Characterization with JWST and the Upcoming Ariel Mission

Model ReleasesDGX agent

arXiv:2606.23766v1 Announce Type: cross Abstract: The detection and atmospheric characterization of exoplanets have entered a new data-intensive era driven by the James Webb Space Telescope and the up

MedPCFM: Improving Medical Point Cloud Completion by Integrating Point Transformers and Flow Matching

ResearchDGX agent

arXiv:2606.24433v1 Announce Type: cross Abstract: Medical point cloud completion is important for anatomical reconstruction and downstream clinical workflows, yet generative modeling in this setting r

Multimedia and Visual Analytics in the Agentic Era

Model ReleasesDGX agent

arXiv:2504.06138v3 Announce Type: replace-cross Abstract: Professional users need tools to help them gain actionable insights from large multimedia collections. Foundation models and AI agents have ra

Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs

Model ReleasesDGX agent

arXiv:2606.23938v1 Announce Type: new Abstract: Driving VLA models incorporating Chain-of-Thought (CoT) reasoning are attractive because they leverage pretrained VLM representations and expose interme

One Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline

Model ReleasesDGX agent

arXiv:2606.23767v1 Announce Type: new Abstract: Headline accuracies on the Tuebingen cause-effect pairs are routinely compared across papers even though each is measured under its authors' own protoco

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought

TutorialsDGX agent

arXiv:2606.24539v1 Announce Type: new Abstract: Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene

REDI-Match: Rotation-Equivariant Distillation for Efficient and Robust Dense Matching

Model ReleasesDGX agent

arXiv:2606.24330v1 Announce Type: new Abstract: Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment

Model ReleasesDGX agent

arXiv:2606.23700v1 Announce Type: cross Abstract: Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates thro

Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation

ResearchDGX agent

arXiv:2507.09839v2 Announce Type: replace Abstract: An increasing number of NLP applications interact with large language models (LLMs) through black-box APIs, making prompt engineering critical for c

T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph

Model ReleasesDGX agent

arXiv:2606.24145v1 Announce Type: new Abstract: Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explici

Tri-Efficient Transfer Learning for Point Cloud Videos

Model ReleasesDGX agent

arXiv:2606.24175v1 Announce Type: new Abstract: While point cloud foundation models have significantly advanced point cloud video understanding, existing parameter-efficient fine-tuning (PEFT) methods

We've provided some updated results on Mistral OCR that make use of the annotation feature for charts. The overall score is ahead of GPT-5.5…

Model ReleasesDGX agent

We've provided some updated results on Mistral OCR that make use of the annotation feature for charts. The overall score is ahead of GPT-5.5 and just behind Gemini 3.1 Pro, which is quite impressive f

When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

Model ReleasesDGX agent

arXiv:2606.23937v1 Announce Type: cross Abstract: Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model. We test t

When Top-1 Fails: Calibrating LoRA Monitors for Masked Diffusion LMs

Model ReleasesDGX agent

arXiv:2606.24119v1 Announce Type: cross Abstract: Discrete diffusion language model (DLM) fine-tuning inherits inexpensive diagnostics from denoising-time confidence monitors, but their PEFT-training

23 Jun 2026

AI Agents Can Already Autonomously Perform Experimental High Energy Physics

Model ReleasesDGX agent

arXiv:2603.20179v3 Announce Type: replace-cross Abstract: Large language model-based AI agents are now able to autonomously execute substantial portions of a high energy physics (HEP) analysis pipelin

AI-Augmented Thyroid Scintigraphy for Robust Classification of Disease

ResearchDGX agent

arXiv:2503.00366v3 Announce Type: replace-cross Abstract: Thyroid scintigraphy is vital for diagnosing thyroid disorders, yet deep learning (DL) models in this domain often struggle with limited, imba

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation

Model ReleasesDGX agent

arXiv:2606.15127v2 Announce Type: replace Abstract: Reasoning models are increasingly used in settings where the final answer is not the only object of review: educational tools may show students inte

Concept Alignment Contrast and Long-Short Prompt Memory for Test-Time Adaptation of SAM3 in Medical Image Segmentation

SafetyDGX agent

arXiv:2606.22963v1 Announce Type: new Abstract: Concept segmentation models like Segment Anything Model 3 (SAM3) show strong generalization on natural images, yet their performance degrades in medical

CoRDE: Concept-Prior Routed Diffusion Experts for Structural Generalization in Robot Manipulation

Model ReleasesDGX agent

arXiv:2606.21935v1 Announce Type: new Abstract: Diffusion models excel at capturing multi-modal action distributions in robot imitation learning. However, in multi-task and long-horizon scenarios, mon

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming

Model ReleasesDGX agent

arXiv:2606.22476v1 Announce Type: new Abstract: Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cr

Data Pruning: Redundant, Problematic, and Interdependent Samples

Model ReleasesDGX agent

arXiv:2606.21916v1 Announce Type: new Abstract: The performance of deep learning models is affected by not only data quantity but also data quality. Data pruning is a process by which practitioners ca

Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers

ResearchDGX agent

arXiv:2606.21678v1 Announce Type: new Abstract: Language models can generate plausible rationales for their predictions, but these explanations may not faithfully represent the model's internal reason

Deep Learning-Based Sign Language Recognition from Videos and Cross-Lingual Translation to Indian Vernaculars

Model ReleasesDGX agent

arXiv:2606.22494v1 Announce Type: cross Abstract: Sign language is a primary mode of communication for the global deaf and hard-of-hearing community, yet automated tools that recognize sign gestures f

Deep Learning for Individual Heterogeneity

SafetyDGX agent

arXiv:2010.14694v4 Announce Type: replace-cross Abstract: This paper integrates deep neural networks (DNNs) into structural models to increase flexibility and capture rich heterogeneity while preservi

Demystifying Numerical Instability in LLM Inference: Achieving Reproducible Inference for Mission-Critical Tasks with HEAL

Model ReleasesDGX agent

arXiv:2606.21023v1 Announce Type: new Abstract: As Large Language Models (LLMs) deploy into mission-critical domains (e.g., finance, medicine, and law), output reproducibility has become a strict syst

DrugBench: Evaluating AI Control Protocols for Medication Harm Mitigation

Model ReleasesDGX agent

arXiv:2606.20663v1 Announce Type: cross Abstract: Large Language Models have the potential to expand and improve the access to clinical information by enabling new ways of interacting with medical kno

Dynamics, stability, and energy efficiency of an energy-recycling rimless wheel with spring-clutch legs

Model ReleasesDGX agent

arXiv:2606.22073v1 Announce Type: new Abstract: This paper proposes an energy-recycling rimless wheel with spring-clutch legs. The proposed mechanism uses a lockable clutch to store part of the impact

Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

SafetyDGX agent

arXiv:2504.05520v4 Announce Type: replace Abstract: Reinforcement finetuning (RFT) has shown great potential for enhancing the mathematical reasoning capabilities of large language models (LLMs), but

Flowing With Purpose: Latent Action Guided Flow Matching Policies For Robotic Manipulation

Model ReleasesDGX agent

arXiv:2606.23420v1 Announce Type: new Abstract: Flow matching has recently become a new standard for behavior cloning in robotic manipulation. However, state-of-the-art flow matching policies suffer f

GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation

Model ReleasesDGX agent

arXiv:2606.23669v1 Announce Type: new Abstract: Text-to-image models can generate visually plausible city streets, but whether their outputs correspond to a requested road segment rather than a generi

GRAG: Generic Response-Augmented Generation Framework for Personalized Conversational Systems

Model ReleasesDGX agent

arXiv:2606.21097v1 Announce Type: cross Abstract: Deploying highly capable personalized conversational agents in resource-constrained or privacy-sensitive environments remains a significant challenge.

In-Context Molecular Property Prediction with LLMs: A Blinding Study on Memorization and Knowledge Conflicts

Model ReleasesDGX agent

arXiv:2603.25857v2 Announce Type: replace Abstract: The capabilities of large language models (LLMs) have expanded beyond natural language processing to scientific prediction tasks, including molecula

Introducing Mistral OCR 4. It creates structure with bounding boxes, block classification, and inline confidence scores in 170 languages. 🧵…

Model ReleasesDGX agent

Mistral OCR 4 is an optical character recognition model that extracts text with structural information including bounding boxes and block classification across 170 languages. The system provides inlin

Is Our Benchmark Enough? An Analysis of Continual Learning for MLLMs

Model ReleasesDGX agent

arXiv:2606.20961v1 Announce Type: new Abstract: Continual adaptation is essential for multimodal large language models (MLLMs) deployed across evolving domains, but the state-of-the-art MR-LoRA method

Learning Bug Context for PyTorch-to-JAX Translation with LLMs

Model ReleasesDGX agent

arXiv:2510.09898v2 Announce Type: replace Abstract: Large language models (LLMs) have shown strong performance on code translation between widely used programming languages. However, translation becom

MoECodec: Image Compression for joint human and machine perception via Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2606.21033v1 Announce Type: cross Abstract: Image compression for machines calls for a unified codec that serves multiple downstream vision tasks. Existing approaches either adopt task-specific

MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning

Model ReleasesDGX agent

arXiv:2606.23061v1 Announce Type: new Abstract: Motion instruction generation in cross-video comparison aims to produce corrective feedback that describes the differences between a query and a referen

Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?

Model ReleasesDGX agent

arXiv:2606.23676v1 Announce Type: new Abstract: AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostly in finite-variance regimes. This

Physics-Guided Dual-Stream Heterogeneous Graph Neural Network for Predicting Full-Field Structural Response of Stiffened Panels

Model ReleasesDGX agent

arXiv:2606.20916v1 Announce Type: new Abstract: Iterative design and optimization of large, complex structures require fast and accurate prediction of stress, displacement, and other fields. Finite el

Priority-Aware Learning-Unlearning Correction for Dynamic Decentralized LoRA Fine-Tuning

Model ReleasesDGX agent

arXiv:2606.22878v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed at the network edge to provide pervasive generative AI services, decentralized federated learn

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations

Model ReleasesDGX agent

arXiv:2606.22766v1 Announce Type: new Abstract: Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing met

Revealing the Pitfalls and Re-Evaluating the Advancement of Heterophilic Graph Learning

Model ReleasesDGX agent

arXiv:2409.05755v3 Announce Type: replace Abstract: Over the past decade, Graph Neural Networks (GNNs) have achieved great success on machine learning tasks with relational data. However, recent studi

RoverDevKit: An open, physics-grounded tradespace toolkit for conceptual design of lunar micro-rovers

Model ReleasesDGX agent

arXiv:2606.21755v1 Announce Type: new Abstract: Pre-Phase-A design of lunar micro-rovers is dominated by tightly coupled mobility, power, thermal, and mass trades, yet conceptual-design tooling for th

S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix

ResearchDGX agent

arXiv:2508.08048v2 Announce Type: replace Abstract: While video generation models excel at producing high-quality monocular videos, generating 3D stereoscopic and spatial videos for immersive applicat

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training

Model ReleasesDGX agent

arXiv:2606.21090v1 Announce Type: cross Abstract: Self-improvement can self-regress. In REINFORCE post-training for code, a model can quickly improve on its optimized metric and then collapse within t

Skill Coverage: A Test Adequacy Metric for Agent Skills

Model ReleasesDGX agent

arXiv:2606.20659v1 Announce Type: cross Abstract: Agent skills encode reusable procedural knowledge that guides large language model agents across tasks and execution contexts. Existing evaluations pr

Subspace-Constrained Federated Learning with Low-Rank Adaptation

Local AiDGX agent

arXiv:2606.22724v1 Announce Type: new Abstract: Federated low-rank adaptation methods are attractive for fine-tuning large models under communication and privacy constraints, but heterogeneous client

Temporally Aware Densification for Dynamic 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2606.23212v1 Announce Type: new Abstract: Despite modeling temporal motion, dynamic 3D Gaussian Splatting (3DGS) methods still inherit a static densification strategy that is ill-suited for dyna

Tensor Train Decomposition-based 3D Implicit Full Waveform Inversion with Multi-scale Structural Similarity

ResearchDGX agent

arXiv:2606.22867v1 Announce Type: cross Abstract: Three-dimensional full waveform inversion (3DFWI) is a powerful technique for reconstructing high-resolution subsurface velocity models. However, its

The Pitfall of Scaling Up: Uncovering and Mitigating Popularity Bias Amplification in Scaling Transformer-based Recommenders

SafetyDGX agent

arXiv:2606.21911v1 Announce Type: cross Abstract: We identify a critical pitfall in scaling transformer-based sequential recommenders: while increasing model size improves recommendation accuracy, it

When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents

Model ReleasesDGX agent

arXiv:2606.22864v1 Announce Type: new Abstract: Hidden-state probing -- a linear classifier on a frozen vision-language model's internal activations -- has emerged as an attractive evaluation tool for

With agentic coding, complexity compounds in a mechanical way: unnecessary code ends up in the codebase, moves to the context window, degrad…

Model ReleasesDGX agent

With agentic coding, complexity compounds in a mechanical way: unnecessary code ends up in the codebase, moves to the context window, degrades the model's reasoning abilities, leads to more unnecessar

22 Jun 2026

PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters

ToolsDGX agent

PP-OCRv6 is a multilingual optical character recognition (OCR) model family released by PaddlePaddle that supports 50 languages across multiple parameter sizes ranging from 1.5M to 34.5M. The model se

← Previous
1…394395396397398…1051
Next →