LFM2.5-2.6B model+KV cache quantization report
LFM2.5-2.6B is a new tiny model by LiquidAI, with benchmarks that put it head to head with much larger models. I've run llama-perplexity on many model GGUF quants, crossed with many KV cache quants, t
Knowledge catalogue
LFM2.5-2.6B is a new tiny model by LiquidAI, with benchmarks that put it head to head with much larger models. I've run llama-perplexity on many model GGUF quants, crossed with many KV cache quants, t
arXiv:2608.05727v1 Announce Type: cross Abstract: Neural Audio Codecs are widely adopted in speech generation and editing. However, existing neural audio codecs are not idempotent: across the paper's
arXiv:2608.05739v1 Announce Type: new Abstract: Mobile image denoising requires both good restoration quality and low computational cost. In addition, it's annoying to collect large-scale LQ-GT clean
Live now: our Local AI Track from AI Engineer World's Fair 2026, brought to you by @nvidia. Thesis: frontier intelligence is becoming something you own. https://www.youtube.com/watch?v=KB41dTlX1Uc&lis
A fresh llama.cpp PR (#26689) changes what looks like a tiny SYCL FlashAttention dispatch decision. With a quantized KV cache ('q4_0' / 'q8_0'), decode was being sent through the VEC kernel. On the au
arXiv:2608.06135v1 Announce Type: new Abstract: Large Language Models (LLMs) such as ChatGPT and Claude are widely used for information retrieval and problem-solving. Recent work has focused on improv
arXiv:2608.05356v1 Announce Type: new Abstract: High-definition 3D LiDAR maps are important for autonomous driving and smart-city services, which require reliable detection of object-level changes in
arXiv:2604.01280v2 Announce Type: replace-cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires Multimodal Large Language Models (MLLMs) to identify and combine fine-grained visu
arXiv:2608.05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain beh
arXiv:2608.05817v1 Announce Type: new Abstract: Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visu
arXiv:2607.16284v2 Announce Type: replace Abstract: Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communicat
arXiv:2608.05872v1 Announce Type: cross Abstract: Standard Large Language Models (LLMs) execute layers sequentially. Dynamic layer routing, i.e. search for a different execution path through layers in
arXiv:2608.05850v1 Announce Type: cross Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, it
**Managing AI Coding Costs at Scale** This article addresses the economic challenges of deploying AI coding tools across large teams or organizations. It explores strategies for tracking resource usag
arXiv:2608.05911v1 Announce Type: new Abstract: This paper presents an end-to-end, IIIF-based pipeline that turns the digitised Armenian press of France into an interactive map of the 20th-century Par
arXiv:2510.03997v2 Announce Type: replace Abstract: Understanding how patients perceive their physicians is essential to improving trust, communication, and satisfaction. Patients increasingly consult
arXiv:2608.05857v1 Announce Type: new Abstract: Retrieval-Augmented Generation systems rely on similarity scores to retrieve relevant content, yet scores are not directly comparable across embedding m
arXiv:2608.05209v1 Announce Type: new Abstract: Constructing reliable online HD maps remains challenging in dynamic urban environments due to moving objects and occlusions. While recent works employ f
arXiv:2608.05243v1 Announce Type: cross Abstract: Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and inter
arXiv:2608.06257v1 Announce Type: new Abstract: Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redunda
arXiv:2606.21140v2 Announce Type: replace-cross Abstract: Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs d
arXiv:2608.05472v1 Announce Type: cross Abstract: Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping
arXiv:2608.05878v1 Announce Type: new Abstract: Text-to-image diffusion transformers learn about objects and scenes by learning to generate them, making them strong candidates for training-free zero-s
arXiv:2608.05152v1 Announce Type: cross Abstract: Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior
arXiv:2608.05624v1 Announce Type: new Abstract: Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This pa
As AI agents gain access to enterprise systems and sensitive data, security teams need greater visibility into their actions. Real-time monitoring and policy enforcement are becoming essential as a pa
arXiv:2608.06253v1 Announce Type: new Abstract: Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed Me
arXiv:2608.06183v1 Announce Type: new Abstract: Microarchitecture design space exploration suffers from expansive search spaces and expensive PPA evaluation, leaving only a small simulation budget for
arXiv:2608.06115v1 Announce Type: new Abstract: Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the l
arXiv:2608.06250v1 Announce Type: cross Abstract: In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient d
arXiv:2607.11785v2 Announce Type: replace Abstract: Indoor robotics research increasingly uses micro-UAV platforms whose airframes, electronics, and control software are open to modification. Off-the-
arXiv:2608.05938v1 Announce Type: new Abstract: Medical images are routinely de-identified---names, dates, and other metadata removed---and then shared for research, teaching, and public benchmarks un
arXiv:2608.05726v1 Announce Type: new Abstract: Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluati
arXiv:2608.06046v1 Announce Type: cross Abstract: AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, train
arXiv:2603.05997v2 Announce Type: replace-cross Abstract: Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time intervals a
arXiv:2603.25406v3 Announce Type: replace Abstract: Vision-Language-Action (VLA) models map visual observations and natural-language instructions to robot actions; however, hierarchical and autoregres
arXiv:2608.05825v1 Announce Type: new Abstract: Human social communication, such as affection and intent, is often conveyed in highly implicit ways, where underlying meanings are expressed through ind
arXiv:2608.05409v1 Announce Type: new Abstract: Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep est
Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5, where I had Claude Fable 5 build a full working game
arXiv:2608.05450v1 Announce Type: new Abstract: Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substanti
arXiv:2608.05196v1 Announce Type: new Abstract: Multiple sclerosis (MS) is diagnosed through clinical assessment, magnetic resonance imaging, laboratory evidence when appropriate, and exclusion of bet
arXiv:2608.05346v1 Announce Type: cross Abstract: Time-sensitive networking (TSN) is increasingly integrated into mobile edge computing (MEC) to support applications with stringent latency requirement
arXiv:2608.05340v1 Announce Type: cross Abstract: Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication for time-s
arXiv:2506.14243v4 Announce Type: replace Abstract: LiDAR-based place recognition is critical for long-term autonomous driving without GPS. Existing handcrafted feature methods face dual limitations.
arXiv:2608.05979v1 Announce Type: new Abstract: Machine learning, and deep networks in particular, are increasingly used to derive higher-level Earth observation (EO) products such as annual land-cove
arXiv:2608.05742v1 Announce Type: cross Abstract: Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. While existing
arXiv:2608.06218v1 Announce Type: cross Abstract: We study Muon, a recently proposed matrix-aware optimization method, in the context of the Stiefel manifold. This manifold consists of matrices with o
I swear AA is not the bipartisan they so claim. An open source mode (Qwen 3.8 max) was number 1 on the agentic index, then they just so happen to launch 'v4.1.1' of their index in which they just adju
arXiv:2603.19229v2 Announce Type: replace-cross Abstract: There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural language i
arXiv:2608.05725v1 Announce Type: new Abstract: Visuotactile sensors reconstruct dense contact geometry from measured surface gradients, but host-based processing increases power consumption and intro
arXiv:2608.05418v1 Announce Type: new Abstract: AI tools are being increasingly adopted in policing in the UK and worldwide. Racial bias is a known and well-documented risk, yet representatives of aff
arXiv:2608.06292v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. H
arXiv:2608.05773v1 Announce Type: new Abstract: A geometry-conditioned, neuro-symbolic closed-loop architecture is proposed for laser powder bed fusion, in which a standards-aligned ontology operates
arXiv:2608.05226v1 Announce Type: new Abstract: Neuron counting and segmentation in microscopy images of neuronal cultures is a routine and time-consuming task in neuroscience research, traditionally
arXiv:2511.14907v2 Announce Type: replace Abstract: Computational pathology holds substantial promise for improving diagnosis and guiding treatment decisions. Recent pathology foundation models enable
arXiv:2608.05684v1 Announce Type: cross Abstract: Nonvisual classification of ground condition based on a multimodal sensing approach was investigated for an amoeba-inspired autonomous walking robot.
OpenAI gave a last-minute presentation at the Black Hat security on Wednesday about 'the Hugging Face Incident' (previously on this blog). The video was published yesterday. It's short and information
arXiv:2608.05989v1 Announce Type: new Abstract: Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning
Hi there! I've messaged the Ollama support team 4 times with no response in 3 weeks. This is getting ridiculous. Does anyone have any recommendations as to how I should seek support? I don't want to i
arXiv:2608.06375v1 Announce Type: new Abstract: Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects