AppAgent: Multimodal Agents as Smartphone Users
arXiv:2312.13771v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper i
Knowledge catalogue
arXiv:2312.13771v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper i
Bloomberg: Apple supplier Luxshare raised ~3.1B in its Hong Kong IPO, selling 383.5M shares at ~8 each, the top of its marketed range, and will start trading on Thursday — Apple Inc. supplier Luxshare
Financial Times: Apple's interest in buying CXMT chips has thrust the Chinese memory chip maker and its relationship with Beijing into the AI supply chain spotlight — Sharp turnaround for state-backed
arXiv:2607.03550v1 Announce Type: new Abstract: Human reasoning often operates through qualitative concepts expressed by linguistic labels such as high, low, expensive, or cheap, whose interpretation
arXiv:2607.03513v1 Announce Type: cross Abstract: We present AquaGen, the first all-atom, explicit solvent, periodic-boundary-condition-aware generative model that produces molecular configurations fr
arXiv:2607.04303v1 Announce Type: new Abstract: Learning-based stereo matching models struggle in underwater environments due to scarce in-domain data and the difficulty of extracting discriminative c
ARC Prize 2026: ARC-AGI-3 Milestone #1 Winners Congratulations to the three winners who open-sourced their top-scoring ARC-AGI-3 solutions: 1. @tufalabs - 1.21%, 25K 2. Reki - .867%, 7.5K 3. Md Boktia
arXiv:2601.07475v2 Announce Type: replace-cross Abstract: The emergence of fine-grained numerical formats like NVFP4 presents new opportunities for efficient Large Language Model (LLM) inference. Howe
arXiv:2602.21534v3 Announce Type: replace Abstract: Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interacti
Reece Rogers / Wired: As part of Meta's Muse Image rollout, Instagram users with public accounts need to opt out to block AI generations of their content by other users — As part of Meta's Muse Image
As the @UN’s Global Dialogue on AI Governance wraps up today, I’ve been encouraged by the discussions surrounding AI and its implications for global collaboration and policymaking. I’m hopeful that we
arXiv:2607.02686v1 Announce Type: new Abstract: Reinforcement learning agents operating under partial observability must act on incomplete information, making them natural candidates for guidance from
Autonomous security company Assail Inc. today launched Sidewinder, a rebuilt version of its Ares offensive security platform that runs on a new 31 billion-parameter model designed to audit its own fin
arXiv:2607.05123v1 Announce Type: new Abstract: Recent advances in large language models and programmatic CAD have significantly improved Text-to-CAD generation for individual parts. However, producti
arXiv:2603.05060v2 Announce Type: replace Abstract: Multi--task learning seeks to improve the generalization error by leveraging the common information shared by multiple related tasks. One challenge
arXiv:2607.04113v1 Announce Type: new Abstract: Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor sigma_{min}, at wh
At UMA, we build and own the full stack, from hardware to software. This allows us to bake safety in at all levels, rather than bolt it on as an afterthought. Because trust is everything when robots s
arXiv:2607.04837v1 Announce Type: new Abstract: Large-scale humanoid motion-tracking controllers are commonly improved by reallocating training effort: difficult motions are sampled more often, isolat
arXiv:2512.18176v2 Announce Type: replace Abstract: Accurate segmentation of anatomical structures in medical images is essential for diagnosis and treatment planning. While recent interactive segment
Fireworks AI is hosting a fireside chat at Raise Summit featuring CEO Lqiao and Matt Evantic from Evantic Capital on the Master Stage at 4:00 PM. The event appears to be a discussion panel or intervie
arXiv:2607.03738v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic information in an evolving context. Pr
arXiv:2607.02563v1 Announce Type: cross Abstract: Diffusion-based text-to-image models can synthesize complex and highly structured visual content, yet the emergence and evolution of semantic structur
arXiv:2607.03073v1 Announce Type: cross Abstract: Criminal identification from surveillance imagery has become a critical research area in intelligent forensic surveillance systems due to the increasi
arXiv:2606.16730v2 Announce Type: replace-cross Abstract: We re-interpret Transformer pretraining as a fast-slow, singularly perturbed flow along depth, with untied weights as its non-autonomous featu
arXiv:2607.04590v1 Announce Type: new Abstract: Pairwise human comparisons are a primary interface through which modern AI systems learn human preferences. RLHF and related alignment pipelines typical
arXiv:2605.11404v2 Announce Type: replace Abstract: Large language models (LLMs) can simulate human-like reasoning and decision-making in individual agents. LLM-powered multi-agent systems (MAS) combi
arXiv:2606.12555v2 Announce Type: replace-cross Abstract: Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a
arXiv:2607.02586v1 Announce Type: new Abstract: Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common fo
arXiv:2603.08387v2 Announce Type: replace Abstract: Micro-expression Action Unit (AU) detection identifies localized AUs from subtle facial muscle activations, providing a foundation for decoding affe
arXiv:2607.04311v1 Announce Type: new Abstract: Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity
Australian Payments Plus, a payments platform, improved its development speed and efficiency by integrating OpenAI's ChatGPT and Codex tools into their workflow. The integration enabled faster code ge
arXiv:2607.04383v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Dete
arXiv:2607.04542v1 Announce Type: cross Abstract: Every LLM agent run re-derives its behavior token by token on a frontier model: brilliant, expensive, slow, and unbounded. We present Auto, a compiler
arXiv:2607.03656v1 Announce Type: cross Abstract: Large Language Models are increasingly used to turn natural-language requirements into code. In access control, that shortcut is dangerous: a generate
arXiv:2607.02771v1 Announce Type: new Abstract: Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training
arXiv:2607.04433v1 Announce Type: cross Abstract: The rapid integration of large language model-based agents into recommender systems has driven a shift from static, ranking-based pipelines toward aut
arXiv:2603.10174v2 Announce Type: replace Abstract: Autonomous underwater vehicles (AUVs) are increasingly used to survey coral reefs, yet efficiently locating specific coral species of interest remai
arXiv:2607.02520v1 Announce Type: cross Abstract: Automated research agents increasingly generate code, retrieve literature, and draft scientific artifacts, but they often fail to verify whether gener
arXiv:2604.03329v2 Announce Type: replace-cross Abstract: Automatic violence detection from video is challenging because violent interactions may be distant, occluded, or only partially visible. Audio
arXiv:2607.02968v1 Announce Type: new Abstract: Massive Activations (MAs) have been widely observed in Transformer-based models, yet their structure and functional roles in Diffusion Transformers (DiT
b9893 is a release build of llama.cpp with Windows OpenVINO 2026.2.1 support . The llama.cpp project uses continuous build-tagged releases rather than traditional semantic versioning , making b9893 on
Release b9894 of llama.cpp was published on July 7, 2026 , and includes a Vulkan backend fix to check src0 type in GGML_OP_SET_ROWS to avoid failures due to unimplemented f16 support . Llama.cpp is th
Release b9895 of llama.cpp includes a fix for speculative inference out-of-bounds read in ngram-map on prompt shrink, with ~2x performance gains in PP_Speed for FP32, Q4_0 and Q8_0 models. The release
Release b9902 of llama.cpp was released on July 7, 2026 , and includes support for SYCL operations including cross_entropy_loss and cross_entropy_loss_back . The release provides builds across multipl
arXiv:2607.03007v1 Announce Type: cross Abstract: Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet these gains oft
Banger compression paper from NVIDIA. (bookmark it) Bigger MoE models keep winning on quality, but serving them at interactive latency is still hard. NVIDIA compresses the hybrid MoE Nemotron-3-Super
arXiv:2607.03981v1 Announce Type: cross Abstract: Memes have become influential communication tools on social media, combining viral visuals with concise messaging to convey impactful ideas. While sub
Barracuda Networks Inc. today announced that it has acquired Evo Security Inc., a startup with a platform for managing employee access to business applications. The terms of the deal were not disclose
arXiv:2607.03891v1 Announce Type: new Abstract: 3D reconstruction of articulated objects from a single image is challenging because large training datasets with paired image and 3D supervision are dif
arXiv:2412.05894v2 Announce Type: replace-cross Abstract: Federated learning (FL) enables collaborative analysis of biomedical data without exchanging sensitive patient-level information, but its perf
arXiv:2510.03798v3 Announce Type: replace Abstract: The batched multi-armed bandit (MAB) problem, where rewards are collected in batches, is pivotal in applications like clinical trials. While prior w
arXiv:2506.22675v4 Announce Type: replace-cross Abstract: Invariant prediction [Peters et al., 2016] analyzes feature/outcome data from multiple environments to identify invariant features - those wit
arXiv:2607.05090v1 Announce Type: new Abstract: Lumbar spine degeneration is a major contributor to chronic low back pain and is routinely assessed on MRI using ordinal grading systems, e.g. normal, m
arXiv:2607.04072v1 Announce Type: cross Abstract: Large language models can generate plausible quantum code, but it is unclear whether they can reliably target the specific software development kit (S
arXiv:2607.02671v1 Announce Type: cross Abstract: Benign overfitting and double descent have come to shape our understanding of generalization in deep learning, establishing that overfitting is not on
arXiv:2607.03453v1 Announce Type: cross Abstract: Inference-time alignment methods, such as Best-of-N, offer a flexible alternative to training-based alignment by using reward models to select high-qu
Microsoft's fifth wave of Xbox layoffs in July 2026 resulted in approximately 3,200 job losses, with Bethesda Game Studios losing at least 35 employees and id Software being hit particularly hard, rep
arXiv:2603.06576v2 Announce Type: replace-cross Abstract: The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic
arXiv:2607.03015v1 Announce Type: new Abstract: Forecasting future events has attracted growing attention as a testbed for general-purpose AI. A natural way to ground this evaluation is let the models
arXiv:2607.03017v1 Announce Type: new Abstract: The aging global population drives demand for assistive robots, yet the safety risks and costs of physical testing make Human-in-the-Loop (HITL) simulat