Software Engineering for and with GUI Agent
arXiv:2608.09278v1 Announce Type: cross Abstract: GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity
Knowledge catalogue
arXiv:2608.09278v1 Announce Type: cross Abstract: GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity
arXiv:2608.08200v1 Announce Type: new Abstract: Communication delay remains a central challenge in telerobotics, where it disrupts visuomotor coordination and reduces task precision. Motion scaling is
arXiv:2608.09138v1 Announce Type: cross Abstract: While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by suboptimal execu
arXiv:2608.09745v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, provi
arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim
arXiv:2608.08326v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, most existing me
arXiv:2608.00143v2 Announce Type: replace-cross Abstract: Automated attack chain generation is critical for modern cybersecurity, yet manual construction fails to scale as adversary behaviors expand.
arXiv:2608.07948v1 Announce Type: new Abstract: VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling comp
arXiv:2608.07808v1 Announce Type: cross Abstract: Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than structured
arXiv:2608.07772v1 Announce Type: new Abstract: We study policy-based reinforcement learning under the mu-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables t
arXiv:2608.00818v2 Announce Type: replace Abstract: The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems scale, thei
arXiv:2512.07901v4 Announce Type: replace-cross Abstract: Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. Rational players do not c
arXiv:2608.07493v1 Announce Type: cross Abstract: Current AI disclaimers often fail to function as intended due to warning habituation and a transparency paradox. As AI-generated information becomes p
arXiv:2608.07980v1 Announce Type: cross Abstract: In recent years, the term voiceprint has regained attention, particularly in technological applications and policy-making contexts, often carrying the
There are no lossless transformations of natural-language text Sophie Alpert shares her 'internal policy on acceptable use of AI writing by engineers'. It's a short read (supporting its own recommenda
arXiv:2608.08309v1 Announce Type: cross Abstract: We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: sema
arXiv:2509.23426v3 Announce Type: replace Abstract: AI scientists are emerging computational systems that serve as collaborative partners in discovery. These systems remain difficult to build because
arXiv:2509.24039v2 Announce Type: replace-cross Abstract: If topography is a fundamental feature of the brain, it should influence both how neurons are arranged in space (i.e. explain brain maps) and
arXiv:2608.09529v1 Announce Type: new Abstract: As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) gen
arXiv:2511.18721v4 Announce Type: replace-cross Abstract: The SmoothLLM defense provides a certification guarantee against jailbreaking attacks, but it relies on a strict 'k-unstable' assumption that
arXiv:2608.09052v1 Announce Type: cross Abstract: Semi-supervised adaptation of vision foundation models (VFMs) commonly freezes the pretrained backbone and updates lightweight modules such as LoRA. H
arXiv:2608.08601v1 Announce Type: new Abstract: To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies focus on broad
arXiv:2604.00997v2 Announce Type: replace Abstract: Reward factorization personalizes large language models (LLMs) by decomposing rewards into shared basis functions and user-specific weights. Yet, ex
arXiv:2607.16097v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is l
arXiv:2608.08359v1 Announce Type: new Abstract: Ordinal regression, also called ordinal classification, is classification of ordinal data, in which the underlying target variable is categorical and co
arXiv:2608.09098v1 Announce Type: new Abstract: End-to-end planning has shown strong promise for autonomous driving, but most existing methods are designed for structured urban roads and generalize po
arXiv:2608.08622v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses u
arXiv:2608.09448v1 Announce Type: cross Abstract: Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains d
arXiv:2608.08558v1 Announce Type: new Abstract: World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generali
arXiv:2608.08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks. Despite this progres
arXiv:2608.09447v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch o
arXiv:2608.08453v1 Announce Type: new Abstract: Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Model (LLM) age
arXiv:2608.07911v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard. This makes expert cache
arXiv:2608.07786v1 Announce Type: new Abstract: Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage pipelines, leading to intricate lineage relationships t
arXiv:2608.09730v1 Announce Type: new Abstract: Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicit
arXiv:2608.08471v1 Announce Type: new Abstract: Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously un-addressed
arXiv:2606.15059v2 Announce Type: replace Abstract: Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on
arXiv:2608.07193v1 Announce Type: cross Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fix
arXiv:2608.07065v1 Announce Type: cross Abstract: Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather than single
arXiv:2608.06609v1 Announce Type: new Abstract: Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testin
arXiv:2608.06996v1 Announce Type: new Abstract: This paper presents a sensor-minimal automated assembly system for bidirectional single-row flat ribbon cable harnesses (FRCHs). Unlike conventional peg
arXiv:2608.07419v1 Announce Type: new Abstract: Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Traditional post-hoc temperature scaling is inherentl
arXiv:2608.06727v1 Announce Type: new Abstract: Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture
arXiv:2608.06940v1 Announce Type: new Abstract: LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the effective inform
arXiv:2608.06559v1 Announce Type: new Abstract: Contextual bandits offer a natural framework for sample-efficient personalization, but practical deployment remains difficult under sparse, biased inter
arXiv:2608.06637v1 Announce Type: cross Abstract: Robust aggregation methods are widely used in federated learning to mitigate the impact of adversarial client behavior. Distance-based aggregation rul
arXiv:2608.06908v1 Announce Type: cross Abstract: We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT). WEAT is a
arXiv:2604.00310v2 Announce Type: replace-cross Abstract: Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Mod
arXiv:2608.06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges
arXiv:2608.06871v1 Announce Type: new Abstract: Complex systems, core objects of study in artificial life, model diverse phenomena through nonlinear, feedback-driven interactions that produce emergent
arXiv:2504.17584v2 Announce Type: replace-cross Abstract: Attention-FC Disaggregated (AFD) LLM inference systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-
arXiv:2608.06532v1 Announce Type: new Abstract: LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritativ
arXiv:2608.06674v1 Announce Type: new Abstract: Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in
arXiv:2607.16999v2 Announce Type: replace-cross Abstract: The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing framew
arXiv:2608.07460v1 Announce Type: cross Abstract: While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively i
arXiv:2608.06582v1 Announce Type: new Abstract: Flow-based generative models can efficiently produce candidate structures for crystal structure prediction (CSP), but their pretrained objectives do not
arXiv:2602.20860v2 Announce Type: replace Abstract: While existing unsupervised domain adaptation (UDA) methods greatly enhance target domain performance in semantic segmentation, they often neglect n
arXiv:2608.06939v1 Announce Type: new Abstract: Adverse weather causes diverse and complex image degradations, severely compromising the reliability of computer vision systems. Existing all-in-one res
arXiv:2608.06536v1 Announce Type: cross Abstract: Nonadiabatic dynamics needs an excited-state gradient and an interstate nonadiabatic coupling matrix element (NACME) at every nuclear geometry, and a
arXiv:2608.07147v1 Announce Type: new Abstract: Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from co