Xiaomi-GUI-0 Technical Report
arXiv:2606.31410v1 Announce Type: new Abstract: Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions s
Knowledge catalogue
arXiv:2606.31410v1 Announce Type: new Abstract: Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions s
arXiv:2606.28531v1 Announce Type: cross Abstract: Automatically generated videos from scientific papers are increasingly used for education and research dissemination. However, existing evaluation met
arXiv:2606.18288v2 Announce Type: replace-cross Abstract: This volume develops a knowledge theory of capital for economies in which productive capacity increasingly resides in software, data, models,
arXiv:2606.28635v1 Announce Type: new Abstract: Inverse rendering requires separating illumination from surface materials, which is highly ambiguous due to their tight coupling in observed images. Whi
arXiv:2603.17975v2 Announce Type: replace Abstract: We present AHOY, a method for reconstructing complete, animatable 3D Gaussian avatars from in-the-wild monocular video despite heavy occlusion. Exis
Orlando Crowcroft / Financial Times: AI and vibe coding fuel a surge in game releases; ATTN Economy says 181K mobile games launched in six months to May, up 118% on iOS and 73% on Android YoY — More i
arXiv:2606.29335v1 Announce Type: cross Abstract: Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch between training
arXiv:2606.29167v1 Announce Type: new Abstract: Finding dense correspondences between 3D shapes is a fundamental yet unresolved challenge, especially in real-world environments. These environments pre
As open models get stronger, more workloads move into the competitive inference market. That pushes the real fight toward speed, cost, reliability, and control. Together AI is where open models become
arXiv:2510.20640v2 Announce Type: replace Abstract: In this paper, we present DiRecGNN, an attention-enhanced entity recommendation framework for monitoring cloud services at Microsoft. We provide ins
arXiv:2606.28345v1 Announce Type: cross Abstract: LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms vary across cultures by age, status,
arXiv:2606.29925v1 Announce Type: new Abstract: As deep learning models are increasingly deployed in high-stakes applications, providing well-calibrated uncertainty estimates has become as critical as
arXiv:2501.16726v2 Announce Type: replace-cross Abstract: Semantic communications aim to enhance transmission efficiency by jointly optimizing source coding, channel coding, and modulation. While prio
arXiv:2409.05980v2 Announce Type: replace-cross Abstract: Rested and Restless Bandits are two well-known bandit settings that are useful to model real-world sequential decision-making problems in whic
arXiv:2606.28946v1 Announce Type: new Abstract: This paper presents a robust Automatic Number Plate Recognition (ANPR) system tailored for Nepali license plates written in Devanagari script. In this p
arXiv:2606.30030v1 Announce Type: new Abstract: Blind image deblurring demands the recovery of high-fidelity details and coherent structures from complex, unknown degradations. Current blind image deb
Common challenge that will come up in the near future is capturing the gains of greater AI intelligence in organizations. High human capital firms need to be set up to benefit from their high-quality
arXiv:2606.30237v1 Announce Type: new Abstract: In our goal to develop personalised dysarthric speech recognition (DSR) models, this study compared the recognition performances of human listeners and
arXiv:2606.30458v1 Announce Type: new Abstract: Text-to-image person re-identification (TIPR) retrieves target persons using natural language descriptions. However, existing methods largely overlook r
arXiv:2606.30293v1 Announce Type: new Abstract: Robotic applications increasingly rely on distributed computational infrastructures that combine embedded devices, edge servers, and cloud resources. Th
arXiv:2606.29907v1 Announce Type: cross Abstract: Cardiac discharge phenotyping informs post-discharge treatment and follow-up, but real-world records are often incomplete and class-imbalanced, increa
arXiv:2606.29314v1 Announce Type: new Abstract: With the growing demand for immersive visual experiences, high-quality omnidirectional images (ODIs) have become increasingly important. However, limita
arXiv:2606.29924v1 Announce Type: new Abstract: Generating 3D hand-object interactions is essential for applications in robotics, XR, and synthetic data generation, where flexible controllability and
arXiv:2506.08319v2 Announce Type: replace-cross Abstract: Uncertainties and disturbances in robotic systems, such as aerodynamic forces, are fundamentally outcomes of physical interactions with the en
arXiv:2606.29094v1 Announce Type: new Abstract: Diffusion language models (DLMs) have recently emerged as a promising alternative to conventional autoregressive language models. By generating multiple
arXiv:2606.30192v1 Announce Type: new Abstract: Sim-to-real transfer remains a major obstacle for reinforcement learning (RL), especially for vision-based control where image observations exacerbate t
Don't fall for Elon Musk's gaslighting. There is a massive, life-or-death difference between well-planned budget cuts and a sudden, reckless halt to promised medical aid. Musk is actively lying about
arXiv:2603.03143v2 Announce Type: replace-cross Abstract: Leveraging the priors of 2D diffusion models for 3D editing has emerged as a promising paradigm. However, multi-view consistency remains chall
arXiv:2606.28592v1 Announce Type: new Abstract: Physical caregiving robots need to assist different users with different tasks in diverse environments, and they come in many embodiments. While substan
arXiv:2504.07660v2 Announce Type: replace Abstract: Facial expression detection requires spotting when expressions occur and recognizing which emotional category they belong to. Despite their close re
arXiv:2606.28449v1 Announce Type: cross Abstract: Background: Digital health technologies allow for frequent, remote gait monitoring in people with multiple sclerosis (MS). However, to differentiate d
arXiv:2606.29384v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable
arXiv:2606.28882v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown the potential to generate code explanations that surpass those of peers in quality, offering promising opportu
arXiv:2606.28404v1 Announce Type: cross Abstract: Artificial intelligence depends on large-scale compute resources and their supporting infrastructure. However, AI governance debates treat compute pri
Reuters: Five Chinese tech and advanced manufacturing companies launched Hong Kong listings today to raise up to 5.6B, led by Apple supplier Luxshare's 3.15B offering — Five Chinese technology and adv
arXiv:2606.30336v1 Announce Type: new Abstract: We introduce FlexTab, a flexible encoder-decoder architecture for in-context learning on tabular data that pairs a single, task-agnostic encoder with a
arXiv:2606.28745v1 Announce Type: new Abstract: Diffusion prior-based methods have shown impressive results in real-world image super-resolution (ISR), yet two key challenges persist: balancing pixel-
arXiv:2511.13216v2 Announce Type: replace Abstract: Deployment of legged robots for navigating challenging terrains (e.g., stairs, slopes, and unstructured environments) has gained increasing preferen
arXiv:2606.29256v1 Announce Type: cross Abstract: In recent years, models based on the Transformer architecture have seen widespread applications and have become one of the core tools in the field of
arXiv:2512.07854v2 Announce Type: replace Abstract: Traffic forecasting task is significant to modern urban management. Recently, there is growing attention on large-scale forecasting, as it better re
arXiv:2606.30179v1 Announce Type: new Abstract: Accurate identification of resistor values from unconstrained images remains a challenging computer vision task due to variations in lighting, orientati
arXiv:2606.28813v1 Announce Type: cross Abstract: Human videos are a scalable source of supervision for robot manipulation, as they are abundant and naturally capture rich object interactions. However
arXiv:2606.30404v1 Announce Type: new Abstract: Understanding and navigating human-centered environments over extended periods of time while considering human behavior and routines remains a fundament
I wrote about how the rapid rise in AI abilities is leading to both a transformation in how AI is used at work, and the sort of sudden lurches in policies and markets we have been seeing in recent wee
if US bans open source because China is leading, then 1. impossible to enforce(you just need to download a zip for weights) 2. Increases the cost of using AI for US companies and consumers, while Chin
arXiv:2606.28623v1 Announce Type: new Abstract: Effective sub-typing (also known as grouping or clustering) of patients using their electronic health record (EHR) data can greatly inform precision med
inference reliability has historically been a tax on devs that only large well-funded startups could afford: reserve GPUs in advance, sign a contract, guess your peak throughput requirements. everyone
Just hardcore engineering This required: - advanced optics solutions to be able to see vessels through an opaque layer of tissue - precision manufacturing to machine micrometer scale features in our n
arXiv:2606.29928v1 Announce Type: cross Abstract: Multimodal Large Models have significantly advanced automated breast ultrasound diagnosis. However, most existing frameworks utilize opaque, end-to-en
Secure business communications startup LeapXpert Inc. said today it has bagged 180 million in a growth round of funding to build out its artificial intelligence capabilities and generate valuable inte
arXiv:2606.28538v1 Announce Type: new Abstract: We investigate domain adaptation of modern BERT models in the legal domain. We further pre-train ModernBERT on all US court opinions using the masked la
arXiv:2606.29437v1 Announce Type: cross Abstract: The growing use of Large Language Models (LLMs) in education, software engineering, academic writing, and technical documentation raises a key questio
arXiv:2606.28329v1 Announce Type: cross Abstract: The growing adoption of AI in healthcare, particularly in preventive care, highlights the critical need for accessibility and precision in Medical Que
arXiv:2606.29844v1 Announce Type: cross Abstract: The quadratic computational cost of traditional attention mechanisms poses a major bottleneck to the scalability and practical deployment of large lan
arXiv:2606.28622v1 Announce Type: new Abstract: Differentiable rendering has emerged as a powerful approach for 3D reconstruction and novel view synthesis. State-of-the-art differentiable rendering me
arXiv:2606.28470v1 Announce Type: new Abstract: We demonstrate how emotional valence influences the order-dependent structure of children's recognition memory: correct recall of a sequence of emotiona
arXiv:2511.19119v2 Announce Type: replace Abstract: Spatial reasoning (SR), the ability to infer 3D spatial information from 2D inputs, is essential for real-world applications such as embodied AI and
arXiv:2606.29237v1 Announce Type: cross Abstract: Robust robot autonomy depends on scene representations that remain stable enough to support localization, navigation, and downstream decision making i
arXiv:2105.09254v4 Announce Type: replace-cross Abstract: In many applications, researchers are interested in the direct and indirect causal effects of a treatment or exposure on an outcome of interes
arXiv:2606.29303v1 Announce Type: new Abstract: We propose a mask-free method for lossless multi-object 3D reconstruction from sparse and occluded real-world views, enabling physically plausible inter