Model Releases
swiss-ai/Apertus-v1.5 70B/8B
https://huggingface.co/swiss-ai/Apertus-v1.5-70B https://huggingface.co/swiss-ai/Apertus-v1.5-8B Apertus 1.5 is a family of 8B and 70B parameter language models designed to advance the state of multil
https://huggingface.co/swiss-ai/Apertus-v1.5-70B https://huggingface.co/swiss-ai/Apertus-v1.5-8B Apertus 1.5 is a family of 8B and 70B parameter language models designed to advance the state of multilingual, multimodal, fully open, and transparent AI. The models support a wide range of languages, handle contexts of up to 262,144 tokens, and it uses only fully open training data whilst delivering performance comparable to other models of similar size. The released models are the result of continued pretraining of Apertus 1.0, adding a multimodal mix of 4T tokens to the 8B model and 2T tokens to the 70B model. Apertus 1.5 thus uses the same architecture as the original release, a decoder-only transformer with the xIELU activation function trained with the AdEMAMix optimizer. Our improved post-training recipe enhances the models' instruction-following and tool-use capabilities and, for the first time, allows developers to enable a thinking mode to improve the models' performance on reasoning tasks. As a first in the Apertus family, the Apertus 1.5 models support multimodal inputs. The model takes images, audio, and text as input and generates text. This enables many new exciting use cases for our developers. Key Features Fully Open Model: Open weights + open data + full training details including all data and training recipes. Massively Multilingual: Supporting a large variety of languages. Responsible Development: Apertus is trained while respecting opt-out consent of data owners (even retroactively) where possible and with methods to prevent memorization of training data. Native Audio & Image Understanding: Apertus 1.5 introduces multimodal support for processing audio and image inputs, enabling more intuitive and versatile interaction beyond text. Reasoning: The models can be switched to thinking mode to reason on the input before generating responses. Long Context: Apertus 1.5 by default supports a context length up to 262,144 tokens, a four-fold increase from our initial Apertus 1.0 release. Improved Instruction-Following: Significant improvements in instruction adherence ensure more predictable and accurate responses to user prompts. Improved Tool Use: Apertus 1.5 has been trained for better tool integration, allowing for more effective use of external tools and APIs. The technical report with further details along with benchmark results, training pipelines, and intermediate checkpoints will be published in the coming weeks. submitted by /u/jacek2023 [link] [comments]
Source: r/LocalLLaMA | 2026-07-24