AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,429
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,935
  • Model Releases23,918
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,429
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,935
  • Model Releases23,918
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

88,429Total entries
1Added by human
88,428Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,648 results
23 Jun 2026

Data Selection Through Iterative Self-Filtering for Vision-Language Settings

ResearchDGX agent

arXiv:2606.23611v1 Announce Type: new Abstract: The availability of large amounts of clean data is paramount to training neural networks. However, at large scales, manual oversight is impractical, res

Distribution-Aware Diffusion-LLM for Robust Ultra-Long-Term Time Series Forecasting

SafetyDGX agent

arXiv:2606.23391v1 Announce Type: new Abstract: Time series forecasting is a fundamental machine learning task. Recent work has explored Large Language Models (LLMs) for this purpose due to their stro

Do Location Encoders Capture Spatial Effects? A GeoShapley Benchmark Across Scales

Model ReleasesDGX agent

arXiv:2606.23453v1 Announce Type: new Abstract: Location encoders transform geographic coordinates into high dimensional embeddings for downstream machine learning, but it is unclear how well these re

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

DR-Mamba: Automatic Inference-Time Domain Adaptation for Document Image Binarization via Sample-Conditioned Detail-Background Suppression

Model ReleasesDGX agent

arXiv:2606.22625v1 Announce Type: new Abstract: Degraded document image binarization is sensitive to domain shifts caused by paper aging, bleed-through, stains, shadows, and uneven illumination, and t

DrivingVoxels: Compositional Sparse Voxel Rasterization for Dynamic Driving Scene Reconstruction

Model ReleasesDGX agent

arXiv:2606.23031v1 Announce Type: new Abstract: Reconstructing dynamic urban scenes remains challenging due to the unbounded nature of driving environments and the presence of multiple dynamic objects

Enabling Cloud-Level Accuracy in Edge AI through IoT Data Preprocessing

Local AiDGX agent

arXiv:2606.22496v1 Announce Type: cross Abstract: Large language models (LLMs) offer a natural-language interface for interpreting Internet of Things (IoT) sensor data in smart environments; however,

Enhancing Creativity in 3D Generative Design via a TRIZ-Inspired Text-to-CAD Framework

AgentsDGX agent

arXiv:2606.21378v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have demonstrated significant potential in supporting engineering design tasks, including computer-aided

ENVS: Environment-Native Verified Search for Long-Horizon GUI Agents

Model ReleasesDGX agent

arXiv:2606.22948v1 Announce Type: cross Abstract: As multimodal agents move from interface understanding to real software control, successful trajectory discovery in live desktop environments becomes

Experimenting with the proposed Cross-Origin Storage API in Transformers.js

ToolsDGX agent

This article explores the Cross-Origin Storage API and its implementation within Transformers.js, a JavaScript library for machine learning models. It likely discusses how this API enables secure cros

Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm

Model ReleasesDGX agent

arXiv:2509.23946v3 Announce Type: replace Abstract: Many LLMs plan before they act, yet planning and execution are often still entangled in one long generation trace, enforced only through prompts, or

Factor-Aware Mixture-of-Experts with Pretrained Encoder for Combinatorial Generalization

Model ReleasesDGX agent

arXiv:2606.21100v1 Announce Type: new Abstract: The integration of pretrained encoders with diffusion policies has become a dominant paradigm for visual robotic manipulation. However, it still struggl

Factored Gossip DiLoCo: Reducing Blocking Communication in DiLoCo

Model ReleasesDGX agent

arXiv:2606.22768v1 Announce Type: new Abstract: To make large-scale distributed training practical outside high-bandwidth datacenters, we must reduce blocking, high-volume synchronization. While DiLoC

Faithful Grounded Visual Reasoning via Learned Proxy-Tokens

ResearchDGX agent

arXiv:2606.23354v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in Visual Question Answering (VQA), yet their 'black-box' nature hinders deplo

FirstPass: Grounding AI Scientific Judgment in Multi-Round Editorial Outcomes

Model ReleasesDGX agent

arXiv:2606.20769v1 Announce Type: cross Abstract: AI systems for peer review fail on three fronts: they train on Computer Science and Machine Learning venues alone, ignore the iterative dialogue that

FLFL: Federated Latent Factor Learning for Private Recovery of Spatio-Temporal Signals

ApplicationsDGX agent

arXiv:2606.23091v1 Announce Type: new Abstract: Wireless sensor network (WSNs) stands out as a burgeoning and promising domain in intelligent sensing. Owing to various factors such as sudden sensor ma

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

SafetyDGX agent

arXiv:2606.20867v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-sho

From Markov to Laplace: How Mamba In-Context Learns Markov Chains

ResearchDGX agent

arXiv:2502.10178v2 Announce Type: replace Abstract: While transformer-based language models have driven the AI revolution thus far, their computational complexity has spurred growing interest in viabl

Full Nonlinear Nonholonomic Dynamics and Motion Analysis of a 3-DoF Underactuated Spherical Rolling Robot

ResearchDGX agent

arXiv:2606.22169v1 Announce Type: new Abstract: This paper presents a full nonlinear constrained dynamic model of MonoRollBot, a novel 3-DoF spherical rolling robot driven by a single motor, a lead-sc

GeoRouteNet: Geometry-Enhanced Non-Autoregressive Neural Solver for the Traveling Salesman Problem

Model ReleasesDGX agent

arXiv:2606.22776v1 Announce Type: new Abstract: The traveling salesman problem (TSP) is a canonical NP-hard combinatorial optimization benchmark that tests the representational capacity and generaliza

HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training

SafetyDGX agent

arXiv:2606.20189v2 Announce Type: replace Abstract: Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data

How Omio is building the future of conversational travel

Model ReleasesDGX agent

Omio, a travel platform, is leveraging OpenAI's technology to enhance its services through conversational AI capabilities. This implementation likely enables customers to search for and book travel op

HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks

Model ReleasesDGX agent

arXiv:2603.19822v2 Announce Type: replace Abstract: Existing UAV vision-language navigation (VLN) benchmarks have enabled language-guided flight, but they largely focus on long, step-wise route descri

IMAGIN-4D: Image-Guided Controllable Interaction Generation

Model ReleasesDGX agent

arXiv:2606.23675v1 Announce Type: new Abstract: Generating human-object interactions (HOI) is central to character animation, robotics, AR/VR, and embodied AI. Recent HOI generation methods synthesize

Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems

SafetyDGX agent

arXiv:2606.21970v1 Announce Type: cross Abstract: Full-duplex spoken dialogue models, such as Moshi, enable natural, low-latency voice conversations. However, they remain limited to the audio modality

Interpretable Kolmogorov-Arnold Network with Feature-Isolated Temporal Attention Mechanism for Electricity Load Forecasting

ResearchDGX agent

arXiv:2606.23425v1 Announce Type: new Abstract: Accurate electricity load forecasting is a crucial prerequisite for stable power system operations. While prevalent deep learning models present competi

Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity

Model ReleasesDGX agent

arXiv:2606.20676v1 Announce Type: new Abstract: MLLM-as-a-Judge is conventionally validated by agreement with human annotations, but this metric is undefined when the human pool is culturally heteroge

Leveraging AutoML for Sustainable Deep Learning: A Multi-Objective HPO Approach on Deep Shift Neural Networks

TutorialsDGX agent

arXiv:2606.23208v1 Announce Type: new Abstract: Deep Learning (DL) has advanced various fields by extracting complex patterns from large datasets. However, the computational demands of DL models pose

LOGOS: LiDAR-Only Gaussian Elevation Splatting for Unified Tiny Obstacle Segmentation

Model ReleasesDGX agent

arXiv:2606.21527v1 Announce Type: cross Abstract: Robust obstacle segmentation is essential for the safety of intelligent robots, where LiDAR-based perception systems play a fundamental role in the ro

Mat-Pref: Verifiable-Reward Training Improves Compositional Reasoning in Inorganic Materials

Model ReleasesDGX agent

arXiv:2606.21830v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has driven rapid progress in mathematical and code reasoning, but when extended to science, existi

Metric-Dependent Annotation Saturation for Learning from Label Distributions

ResearchDGX agent

When annotators disagree on a label, the disagreement itself carries signal—and the number of annotators needed to capture it depends on the evaluation metric. We fine-tune NLI models on label distrib

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents

Local AiDGX agent

arXiv:2606.20717v1 Announce Type: new Abstract: Multimodal Large Language Model (MLLM)-based web agents provide practical, high-precision solutions for visual browser automation; however, they inheren

Mixture-of-Experts Graph Transformers for Interpretable Particle Collision Detection

ResearchDGX agent

arXiv:2501.03432v3 Announce Type: replace Abstract: The Large Hadron Collider at CERN produces immense volumes of complex data from high-energy particle collisions, demanding sophisticated analytical

Morphology-Aware Multimodal Representation Learning for Insect Phylogenetic Reconstruction

Model ReleasesDGX agent

arXiv:2606.22077v1 Announce Type: new Abstract: Morphological traits provide important evidence for phylogenetic reconstruction and evolutionary relationship analysis. Recent image-based approaches ha

MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learninga

ResearchDGX agent

arXiv:2606.22220v1 Announce Type: new Abstract: Memorization in machine learning models enables high performance on rare in-distribution samples by capturing their atypical patterns. However, it also

Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings

ApplicationsDGX agent

arXiv:2507.07532v4 Announce Type: replace Abstract: While Prover-Verifier Games (PVGs) offer a promising path toward verifiability in nonlinear classification models, they have not yet been applied to

New Smooth Loss functions for Robust Regression that Closely Approximate Absolute Error and Provide Improved Performance on Datasets With Significant Outliers

ResearchDGX agent

arXiv:2606.22068v1 Announce Type: new Abstract: The performance of supervised machine learning models is directly related to the quality of the training dataset. In particular, the presence of signifi

Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense

Model ReleasesDGX agent

arXiv:2602.09012v2 Announce Type: replace Abstract: The rapid evolution of GUI-enabled agents has rendered traditional CAPTCHAs obsolete. While previous benchmarks like OpenCaptchaWorld established a

OFMU: Optimization-Driven Framework for Machine Unlearning

SafetyDGX agent

arXiv:2509.22483v2 Announce Type: replace Abstract: Large language models deployed in sensitive applications increasingly require the ability to unlearn specific knowledge, such as user requests, copy

OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

AgentsDGX agent

arXiv:2606.22617v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-worl

On the Memorization Behavior of LLMs in Generative Recommendation: Observations, Implications, and Training Strategies

TutorialsDGX agent

arXiv:2606.17276v3 Announce Type: replace-cross Abstract: Generative recommendation (GR) has emerged as a promising direction for recommender systems. Recently, large language models (LLMs) have been

One Size does not Fit All: Heterogeneous Latent Space Alignment for Unsupervised Domain Adaptation

SafetyDGX agent

arXiv:2606.21415v1 Announce Type: new Abstract: Domain shift remains a major obstacle to the reliable deployment of machine learning models in high-stakes environments such as healthcare. While Domain

Orthogonal Discrepancy Kernels for Learning with Partial Physics

Model ReleasesDGX agent

arXiv:2606.21199v1 Announce Type: cross Abstract: We introduce a semi-parametric framework for nonlinear system identification, which decouples discrepancy functions from physics-based components. Ort

OVIG: Optimistic Verification of AI Training Integrity via Gradient Signals

Model ReleasesDGX agent

arXiv:2606.21045v1 Announce Type: cross Abstract: The rapid growth of AI has increased the demand for domain-specific post-training, while the cost and specialization of accelerator infrastructure pus

PHOEBI: An Open-World Benchmark for Bacterial Identification in Phase-Contrast Microscopy

Model ReleasesDGX agent

arXiv:2606.22890v1 Announce Type: new Abstract: Optical microscopy enables rapid, label-free imaging of live bacteria and is the standard instrument for species identification across clinical, environ

Physics-Guided Fully Convolutional Spatiotemporal Learning Toward Digital-Twin-Enabled Microstructure Evolution Prediction

ResearchDGX agent

arXiv:2606.20983v1 Announce Type: new Abstract: Understanding and predicting microstructure evolution is central to materials design, yet purely data-driven spatiotemporal learning models often suffer

Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2606.23604v1 Announce Type: new Abstract: The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion estimation. How

Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning

Model ReleasesDGX agent

arXiv:2606.21447v1 Announce Type: cross Abstract: Automated radiology report generation (RRG) has gained increasing attention because it can reduce the heavy workload of clinical report writing. Howev

Rebuttals Move Peer-Review Scores, but Initial-Review Structure Bounds the Movement

Model ReleasesDGX agent

arXiv:2606.22166v1 Announce Type: cross Abstract: Author rebuttals are the main post-submission window in peer review, but their effect on reviewer scores remains hard to measure because score updates

Robust Zero-Shot Generalization for Open-Vocabulary Action Recognition via Task Arithmetic

ApplicationsDGX agent

arXiv:2606.20734v1 Announce Type: new Abstract: Open Vocabulary Action Recognition (OVAR) enables the recognition of novel actions by leveraging vision-language representations, overcoming the limitat

Robustness Cannot be Reduced to Regularization: Studying Adversarial Training Beyond the Linear Case

ResearchDGX agent

arXiv:2606.21488v1 Announce Type: new Abstract: The vulnerability of ML models to adversarial examples has recently emerged as a major concern. While adversarial training is one of the most effective

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

Model ReleasesDGX agent

arXiv:2606.23344v1 Announce Type: new Abstract: Accurate document layout analysis remains a critical bottleneck for document parsing systems, due to the intricate coupling among heterogeneous document

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

SafetyDGX agent

arXiv:2511.20651v2 Announce Type: replace Abstract: Reinforcement learning (RL) has recently emerged as a promising approach for aligning text-to-image generative models with human preferences. A key

Scaling Diverse Language Generation for 3D Visual Grounding

Local AiDGX agent

arXiv:2606.20946v1 Announce Type: cross Abstract: Developing robust models for 3D visual grounding (3DVG), the localization of entities in a 3D scene described in natural language, is important for en

SCENIC: Semantic-Conditioned Edge-Aware Neural Framework for Structured IoT Command Generation

HardwareDGX agent

arXiv:2606.22296v1 Announce Type: new Abstract: Edge Internet of Things (IoT) agents are often constrained by memory capacity, privacy requirements, communication latency, and recurring inference cost

Streaming Dense Voxel Representations for 3D Occupancy Prediction

Model ReleasesDGX agent

arXiv:2503.22087v3 Announce Type: replace Abstract: In this paper, we explore dense voxel streaming for accurate and efficient 3D occupancy prediction. While dense voxel representations offer fine-gra

Subject-Level Unknown-Identity Identification from Leap Motion Controller 2 Hand Landmarks

Model ReleasesDGX agent

arXiv:2606.22986v1 Announce Type: new Abstract: This work studies subject recognition from Leap Motion Controller 2 (LMC2) hand landmark data under a subject-level unknown-identity identification prot

Surgical Anatomy Recognition with Context Learning using Foundation Representations

ResearchDGX agent

arXiv:2606.22124v1 Announce Type: new Abstract: Accurate recognition of anatomical structures is essential for safe and effective minimally invasive surgery (MIS), yet it remains underexplored in surg

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition

ResearchDGX agent

arXiv:2606.21607v1 Announce Type: new Abstract: Vision-language models such as CLIP have recently achieved strong performance on a wide range of visual understanding tasks. However, most existing mode

Task-Differentiated Atomic Skill Expansion and Routing for Continual Learning Across Highly Heterogeneous Tasks

Model ReleasesDGX agent

arXiv:2606.21307v1 Announce Type: new Abstract: Continual learning (CL) is commonly studied under the assumption that sequential tasks are semantically related or structurally similar. However, in hig

Thalia: A Global, Multi-Modal Dataset for Volcanic Activity Monitoring

Model ReleasesDGX agent

arXiv:2505.17782v4 Announce Type: replace Abstract: Monitoring volcanic activity is of paramount importance to safeguarding lives, infrastructure, and ecosystems. However, only a small fraction of kno

← Previous
1…532533534535536…1061
Next →