Training nGPT
arXiv:2608.01284v1 Announce Type: new Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the
Knowledge catalogue
arXiv:2608.01284v1 Announce Type: new Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the
arXiv:2602.03980v2 Announce Type: replace Abstract: Any language model must decide what to say in novel contexts based on information from similar contexts. But what about contexts that are not novel
arXiv:2608.02320v1 Announce Type: cross Abstract: Traversability analysis is a fundamental capability for autonomous mobile robots operating in unstructured environments. While modern machine learning
arXiv:2608.00042v1 Announce Type: new Abstract: Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-st
arXiv:2608.02238v1 Announce Type: cross Abstract: Ensuring trust in AI systems is essential for the safe and ethical integration of machine learning systems into high-stakes domains such as digital he
arXiv:2608.02270v1 Announce Type: new Abstract: Agriculture 4.0 robotic systems improve field efficiency yet remain too capital-intensive for the fragmented smallholdings that dominate global agricult
arXiv:2608.01733v1 Announce Type: new Abstract: Recent advances in robot learning for manipulation have increased the importance of collecting real-world demonstration data. However, existing robotic
arXiv:2608.01573v1 Announce Type: new Abstract: Recent Vision-Language-Action (VLA) models achieve promising performance in robotic manipulation, typically measured by success rates aggregated over pr
arXiv:2608.00986v1 Announce Type: new Abstract: Current advancements in Multimodal Anomaly Detection (MAD) are largely driven by enhancing multimodal fusion, particularly through the integration of RG
arXiv:2608.00915v1 Announce Type: new Abstract: Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on
arXiv:2608.01220v1 Announce Type: cross Abstract: In recent years, the prevalence of large-scale data-sets and the demand for sophisti-cated learning models have necessitated the development of effici
arXiv:2608.00747v1 Announce Type: new Abstract: Large language models are increasingly integrated into autonomous robotic systems for task planning and control, but this integration exposes them to pr
Horiyuki 'horiyuki42' joined Sakana AI Labs in August 2026 as an Applied Research Engineer Intern. He plans to work full‑time during the university summer recess while concentrating on large‑language‑
arXiv:2607.08918v2 Announce Type: replace Abstract: Cyber-physical power systems are vulnerable to cascading failures caused by interdependencies between power and communication infrastructures. Becau
arXiv:2606.20324v2 Announce Type: replace-cross Abstract: Virtual training environments are software-intensive systems in which reinforcement learning (RL) agents learn, adapt, and demonstrate meaning
arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly i
arXiv:2607.28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across docu
arXiv:2607.29343v1 Announce Type: new Abstract: Artificial intelligence is increasingly embedded in everyday software, making its integration into mobile apps inevitable. However, AI mobile app develo
arXiv:2607.29293v1 Announce Type: new Abstract: Accurate fault location is critical for distribution network reliability. However, increasing distributed energy resource (DER) penetration complicates
arXiv:2607.29440v1 Announce Type: new Abstract: Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions.
arXiv:2607.29302v1 Announce Type: cross Abstract: Reliable robot learning requires a world simulator that can predict action consequences before execution on physical hardware, including risky and fai
arXiv:2603.03322v2 Announce Type: replace-cross Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery. However, rig
arXiv:2502.19135v2 Announce Type: replace Abstract: We present PLANTOR, a framework for generating and executing multi-robot task plans from natural-language task descriptions through LLM-assisted kno
arXiv:2405.07708v3 Announce Type: replace Abstract: Decentralized learning (DL) enables participants to collaboratively train models without a central server, yet it faces significant scalability chal
arXiv:2607.29237v1 Announce Type: new Abstract: LiDAR scene flow estimation has settled into a monoculture: nearly all recent methods share the same feed-forward architecture and the same family of se
arXiv:2607.29078v1 Announce Type: new Abstract: On-policy distillation (OPD) trains student models on their own rollouts to reduce exposure bias. However, in multi-turn agent scenarios, early student
arXiv:2607.29657v1 Announce Type: new Abstract: Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency, and maintenance effectiveness. How
arXiv:2607.29687v1 Announce Type: new Abstract: Sequential robot manipulation requires policies to execute novel combinations of familiar instruction components. However, collecting demonstrations for
arXiv:2603.13415v2 Announce Type: replace Abstract: Valence-arousal (VA) estimation is crucial for capturing the nuanced nature of human emotions in naturalistic environments. While pre-trained vision
arXiv:2607.28750v1 Announce Type: cross Abstract: As mobile applications grow in complexity, traditional End-to-End (E2E) testing frameworks struggle with UI volatility, maintenance overhead, and cros
arXiv:2607.29419v1 Announce Type: cross Abstract: In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guid
arXiv:2607.29383v1 Announce Type: new Abstract: In recent years, with the development of big data technology, increasingly more companies use HDFS for data processing and storage. As a result, the mai
arXiv:2607.29596v1 Announce Type: cross Abstract: Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized so
arXiv:2607.28881v1 Announce Type: new Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common
arXiv:2607.29044v1 Announce Type: new Abstract: Inline notes and collected commentaries are important forms of scholarly communication that evolved within the Confucian exegetical tradition, yet have
arXiv:2607.29019v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances large language models by incorporating external knowledge, but existing pipelines typically operate on p
I tested them by sending the 6 documents, each meant to represent a different document type, through my own webapp and comparing every output against the source. All ran on the same L4 GPU. The docume
arXiv:2309.02332v3 Announce Type: replace-cross Abstract: In the mammalian central nervous system, neurons are organized into populations communicating by spike trains propagating along axonal bundles
It is so good! I don't know why there aren't more people talking about it. Fewer tokens, faster and more accurate than Qwen 3.6 35b a3b. On my setup it's nearly as good as 27b, but 5x faster. And it c
arXiv:2607.28970v1 Announce Type: new Abstract: Hyperspectral image classification is complicated by mixed pixels, spectral ambiguity, class imbalance, and limited annotations. Most current classifier
arXiv:2605.05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographi
arXiv:2607.28684v1 Announce Type: new Abstract: Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult t
arXiv:2607.29473v1 Announce Type: new Abstract: The deployment of deep neural networks for visual affordance segmentation on wearable robots poses may prove critical, due to some conflicting aspects o
arXiv:2607.29228v1 Announce Type: cross Abstract: Swarm and evolutionary algorithms are usually analyzed as complete procedural systems in which nonlinear selection, replacement, and adaptation obscur
arXiv:2607.28683v1 Announce Type: cross Abstract: Large language models benefit from elements in natural language, such as metaphors and analogies in training data and inference input to achieve gener
arXiv:2607.29218v1 Announce Type: new Abstract: With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unfortunately, m
arXiv:2607.28679v1 Announce Type: new Abstract: Multi-agent planning problems arise in a variety of engineering applications, such as multi-robot wildfire fighting and unmanned aerial inspection in fa
arXiv:2512.03438v3 Announce Type: replace Abstract: Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally opt
arXiv:2407.02025v5 Announce Type: replace-cross Abstract: Motivated by applications in chemistry and other sciences, we study the expressive power of message-passing neural networks for geometric grap
arXiv:2607.29398v1 Announce Type: new Abstract: Diffusion models have revolutionized generative tasks but incur high latency due to iterative denoising. While cache-based strategies accelerate inferen
arXiv:2607.28877v1 Announce Type: cross Abstract: Verification consumes the majority of modern chip design effort, yet the formal verification tools that provide mathematical guarantees of correctness
arXiv:2607.28629v1 Announce Type: new Abstract: The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural und
arXiv:2602.23670v2 Announce Type: replace Abstract: Pneumatic artificial muscles (PAMs) enable compliant actuation for soft wearable, assistive, and interactive robots. When arranged antagonistically,
arXiv:2607.29156v1 Announce Type: new Abstract: AI-generated image forgeries are becoming increasingly realistic and difficult to characterize with fixed manipulation patterns. As generative models co
arXiv:2607.29555v1 Announce Type: new Abstract: Lacoste-Julien and Jaggi conjectured in 2015 that the pyramidal width of a polytope cannot increase when a vertex is added, provided that every old poin
An exclusive report by Reuters today has surfaced evidence that suggests Chinese artificial intelligence firms have been leveraging the outputs of American frontier models developed by OpenAI Group PB
arXiv:2607.28776v1 Announce Type: new Abstract: Generative machine learning is increasingly used for inorganic crystal structure generation. Most models and the corresponding evaluation approaches rel
arXiv:2607.29033v1 Announce Type: new Abstract: Existing methods for adapting 2D foundation models such as SAM to 3D volumes either process slices independently---ignoring inter-slice context---or req
arXiv:2209.05838v2 Announce Type: replace Abstract: Visual layouts of graphs representing SAT instances can highlight the community structure of SAT instances. The community structure of SAT instances
arXiv:2607.28990v1 Announce Type: new Abstract: Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution envi