b9124
B9124 is a build release from the llama.cpp project, created on May 12, 2026 . Llama.cpp is an open source software library that performs inference on various large language models such as Llama, deve
Knowledge catalogue
B9124 is a build release from the llama.cpp project, created on May 12, 2026 . Llama.cpp is an open source software library that performs inference on various large language models such as Llama, deve
arXiv:2605.08136v1 Announce Type: cross Abstract: Visual perception plays a central role in competitive robotics, where environmental variations can directly affect real-time detection performance. Th
arXiv:2605.10223v1 Announce Type: new Abstract: Current large language model agent frameworks prioritize autonomy but lack the governability mechanisms required for enterprise deployment. High-risk wr
arXiv:2505.24859v3 Announce Type: replace-cross Abstract: Steering vectors are a lightweight method for controlling text properties by adding a learned bias to language model activations at inference
arXiv:2605.10195v1 Announce Type: new Abstract: Tree-of-Thought (ToT) reasoning structures Large Language Model (LLM) inference as a tree-based search, demonstrating strong potential for solving compl
arXiv:2511.21104v3 Announce Type: replace Abstract: Large language models can generate plausible code, but remain brittle for formal verification in proof assistants such as Lean. A central scalabilit
arXiv:2605.08862v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has become a cornerstone for improving the performance of Large Language Models (LLMs). However, its rollout phase constit
arXiv:2605.10598v1 Announce Type: new Abstract: Large language models (LLMs) have emerged as powerful tools for automatic algorithm design (AAD). However, existing pipelines remain inefficient. They o
arXiv:2605.08965v1 Announce Type: new Abstract: Despite strong performance of Multimodal Large Language Models (MLLMs) on multimodal tasks, predicting whether and why an image is persuasive remains ch
arXiv:2605.08938v1 Announce Type: new Abstract: Fourier Neural Operators (FNOs) can greatly accelerate PDE simulation, but they are often used without formal guarantees that they preserve basic physic
arXiv:2605.09795v1 Announce Type: new Abstract: This paper presents our systems and results for the Hope Speech Detection in Code-Mixed Tulu Language shared task at the Sixth Workshop on Speech, Visio
arXiv:2603.02678v3 Announce Type: replace Abstract: This paper argues for recognizing an emerging paradigm of causal learning by wisdom of the crowd. Recent developments in government, industry, and r
Cluster magicians and GPU whisperers, come join us! We’re looking for supercomputing engineers to build the infrastructure behind real-time interactive models, Tinker, and large-scale training: schedu
arXiv:2511.09493v2 Announce Type: replace Abstract: Motivated by undetectable risks in generative AI, we study a general robust aggregation problem: how to aggregate several probability distributions
arXiv:2605.09869v1 Announce Type: cross Abstract: Zero-shot object navigation has advanced rapidly with open-vocabulary detectors, image--text models, and language-guided exploration. However, even af
arXiv:2605.09085v1 Announce Type: new Abstract: Density estimation is a central primitive in probabilistic modeling, yet continuous, discrete, and mixed-variable domains are often treated by separate
arXiv:2605.08561v1 Announce Type: cross Abstract: Density estimation and reliable prediction regions for outputs are crucial in supervised and unsupervised learning. While conformal prediction effecti
arXiv:2510.06637v3 Announce Type: replace-cross Abstract: Despite advances in test-time scaling and diffusion finetuning, guidance for Auto-Regressive Diffusion Models (ARDMs) remains underexplored. W
arXiv:2605.08856v1 Announce Type: new Abstract: Autoregressive neural simulators now match classical solvers on short-horizon prediction of physical systems, yet their accuracy degrades rapidly when r
arXiv:2605.09253v1 Announce Type: cross Abstract: While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionately drives
arXiv:2602.05243v2 Announce Type: replace-cross Abstract: Transformers achieve strong accuracy but incur high compute and memory cost. Structured pruning reduces inference cost, but most methods rely
arXiv:2511.22565v2 Announce Type: replace Abstract: Neural methods for Complex Query Answering (CQA) over knowledge graphs (KGs) are widely believed to learn patterns that generalize beyond explicit g
arXiv:2605.09414v1 Announce Type: new Abstract: Emojis are widely used in online financial communication, but it is unclear whether they provide transferable sentiment signals across languages, platfo
arXiv:2605.09242v1 Announce Type: cross Abstract: Automated grading of diabetic retinopathy (DR) faces several critical challenges: subtle inter-grade visual distinctions in fine-grained lesion patter
arXiv:2605.09802v1 Announce Type: cross Abstract: Vision-language models (VLMs) enable text-guided object detection but degrade severely under cross-view scenarios where ground and aerial viewpoints d
arXiv:2605.10688v1 Announce Type: new Abstract: Event identification in continuous neural recordings is a critical task in neuroscience. Decoding in EEG is dominated by classifying windows aligned to
arXiv:2503.18273v3 Announce Type: replace Abstract: In recent years, Islamophobia has gained significant traction across Western societies, fueled by the rise of digital communication networks. This p
arXiv:2605.09065v1 Announce Type: new Abstract: Scene graphs (SGs) represent objects and their relationships as structured graphs, enabling applications in image generation, robotics, and 3D understan
arXiv:2605.10647v1 Announce Type: new Abstract: Trajectories are nowadays valuable information for a wide range of applications. However they are also inherently sensitive, as they contain highly pers
arXiv:2605.09331v1 Announce Type: new Abstract: Modern Large Language Model (LLM) training is fundamentally bottlenecked by pathologically flat saddle points in extreme high-dimensional landscapes. Mo
arXiv:2605.09697v1 Announce Type: new Abstract: In many real-world computer vision applications, including medical imaging and industrial inspection, binary classification tasks are characterized by a
arXiv:2605.10765v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, yet real-world deployment often requires continual cap
arXiv:2605.10923v1 Announce Type: cross Abstract: Large language model agents increasingly rely on external skills to solve complex tasks, where skills act as modular units that extend their capabilit
arXiv:2605.08723v1 Announce Type: new Abstract: Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using onl
arXiv:2510.19414v2 Announce Type: replace-cross Abstract: The growing prevalence of speech deepfakes has raised serious concerns, particularly in real-world scenarios such as telephone fraud and ident
arXiv:2605.08273v1 Announce Type: cross Abstract: Accurate traffic prediction is essential for optimizing transportation systems, enhancing resource allocation, and improving overall urban administrat
arXiv:2605.08769v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems have shown strong potential on complex tasks through agent specialization, tool use, and collaborat
arXiv:2605.09777v1 Announce Type: cross Abstract: Gradient-based preference optimization methods for large language model (LLM) alignment suffer from preference collapse, converging to narrow behavior
arXiv:2507.11185v2 Announce Type: replace-cross Abstract: Heart disease continues to pose a critical worldwide health issue, more specifically in areas with insufficient access to healthcare infrastru
arXiv:2605.10604v1 Announce Type: cross Abstract: Designing fair algorithmic decision systems requires balancing model performance with fairness toward affected individuals: More fairness might requir
arXiv:2605.10230v1 Announce Type: new Abstract: Molecular optimization seeks to improve a molecule through small structural edits while preserving similarity to the starting compound. Recent language-
arXiv:2605.10474v1 Announce Type: cross Abstract: Analog neural networks are gaining attention due to their efficiency in terms of power consumption and processing speed. However, since analog neural
arXiv:2605.10351v1 Announce Type: new Abstract: Reliable inference requires that artificial intelligence (AI) models provide trustworthy uncertainty estimates, not merely accurate predictions. Recent
arXiv:2605.08324v1 Announce Type: cross Abstract: Diabetic Retinopathy (DR) is a common complication of diabetes that can lead to blindness of people. Detecting DR at the earliest stage is essential t
arXiv:2605.08767v1 Announce Type: new Abstract: Recent advances in generative modeling have enabled significant progress in structure-based drug design (SBDD). Existing methods typically condition mol
arXiv:2605.08819v1 Announce Type: new Abstract: Deep learning techniques have revolutionised medical imaging, improving diagnostic accuracy and enabling both more accurate and earlier disease detectio
arXiv:2405.09570v2 Announce Type: replace-cross Abstract: Heart murmurs are abnormal sounds caused by turbulent blood flow in the heart. Several diagnostic methods are available to detect heart murmur
arXiv:2603.09007v2 Announce Type: replace-cross Abstract: Audio deepfake detection aims to detect real human voices from those generated by Artificial Intelligence (AI) and has emerged as a significan
arXiv:2603.14681v2 Announce Type: replace Abstract: Bayesian change-point and segmentation models provide uncertainty-aware piecewise-constant representations of ordered data, but exact inference is o
arXiv:2605.09984v1 Announce Type: cross Abstract: Recent 4D generation methods complete scene-level missing information using generative models and reconstruct the scene into radiance-based representa
arXiv:2605.10885v1 Announce Type: new Abstract: Cross-domain few-shot medical image segmentation (CD-FSMIS) requires a model to generalise simultaneously to novel anatomical categories and unseen imag
arXiv:2605.10685v1 Announce Type: new Abstract: Mathematical formulas serve as a language through which humans communicate with nature. Discovering mathematical laws from scientific data to describe n
arXiv:2605.08303v1 Announce Type: cross Abstract: Accurate prediction of structural displacements under external loading is fundamental to structural health monitoring and seismic safety assessment. A
arXiv:2605.10582v1 Announce Type: cross Abstract: This paper proposes a guaranteed defense method for large language models (LLMs) to safeguard against jailbreaking attacks. Drawing inspiration from t
arXiv:2605.09942v1 Announce Type: new Abstract: Memory retrieval in agentic large language model (LLM) systems is often treated as a static lookup problem, relying on flat vector search or fixed binar
arXiv:2605.08210v1 Announce Type: new Abstract: Multi-rater medical image segmentation captures the inherent ambiguity of clinical interpretation, where diagnostic boundaries vary across experts and i
This OpenAI Academy article explores how finance teams leverage Codex, OpenAI's code generation model, to automate financial analysis, reporting, and data processing tasks. It likely covers practical
https://x.com/chhillee/status/2053940218747842619?s=46 In modern ML accelerators, FLOPS have absolutely exploded. Often though, the bottleneck is not FLOPS but memory bandwidth. Similarly, model intel
arXiv:2605.08158v1 Announce Type: cross Abstract: Long-video understanding with multimodal language models suffers from three compounding bottlenecks: heavy decode cost to obtain dense RGB frames, qua
arXiv:2605.10164v1 Announce Type: new Abstract: Dense Associative Memory (DenseAM) is a promising family of AI architectures that is represented by a neural network performing temporal dynamics on an