b10456
sycl: fix thread/block count in quantized cpy kernel launches (#27160) Adjusts the thread/block count to be proportional to the size of the quant, reducing under/over subscription. Largest perf improv
Knowledge catalogue
sycl: fix thread/block count in quantized cpy kernel launches (#27160) Adjusts the thread/block count to be proportional to the size of the quant, reducing under/over subscription. Largest perf improv
ci : push release tag explicitly in release.yml (#27261) Add a 'Create and push git tag' step to the release job, right before the 'Create release' step. The tag is created with git tag and pushed wit
cuda : skip UMA override for HIP builds (#27083) AMD APUs report accurate memory via hipMemGetInfo. Using MemAvailable over-promises on small-carveout systems. fixes #18159 Website: https://llama.app
arXiv:2608.14161v1 Announce Type: new Abstract: LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications. While prior work has mad
arXiv:2608.13990v1 Announce Type: new Abstract: Driven by the attention economy, short-video Recommender Systems (RSs) are primarily optimized to maximize user engagement by promoting videos that capt
arXiv:2608.14403v1 Announce Type: new Abstract: Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foun
Guys, could someone smarter than me explain what makes Deepseek Harness so efficient? I run it with local Qwen 3.8 (Q6). I tried Opencode/Openchamber (my favourite so far), Pi agent and Hermes. New Qw
arXiv:2608.14355v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables the simultaneous profiling of gene expression and tissue morphology, creating an opportunity to learn multimodal re
arXiv:2608.14406v1 Announce Type: new Abstract: Textile-integrated capacitive sensors offer flexible and conformable tactile sensing for wearable electronics and human-robot interaction; however, the
arXiv:2608.14049v1 Announce Type: new Abstract: Robotic manipulation of flat objects is challenging due to the ungraspable configurations and strong variations in object geometry and material. Existin
Gemini and Pixel have entered long‑term partnership agreements with Arsenal FC, FC Barcelona, FC Bayern München, Liverpool FC, and Paris Saint‑Germain to deliver AI‑powered personal intelligence exper
arXiv:2608.14528v1 Announce Type: new Abstract: This study investigates the methodological and theoretical properties of session handover in applications that use large language models. A task may con
arXiv:2503.06821v2 Announce Type: replace Abstract: The exploration of Bird's-Eye View (BEV) mapping technology has driven significant innovation in visual perception technology for autonomous driving
arXiv:2608.13922v1 Announce Type: new Abstract: Detecting distributional changes in high dimension is difficult when neither the pre-change nor post-change density is parametrically specified. We intr
I created 3 SVG prompts, each one rather hard. Perspective: A animated drone view perspective on a park. Beauty: A beach scene with an evil cat Composition: An AGI breaking out of a virtual sandbox pr
I'm back. Yesterday I released the first version of hyper-optimized Qwen3.8-27B inference engine for a RTX 3090, reaching 82 tps on single request and 672 peak. Over the last 24 hours I've been explor
arXiv:2512.14338v4 Announce Type: replace Abstract: Many learning problems are organized by group symmetries. While invariance is often imposed through architectures or group averaging, we ask when it
arXiv:2608.14457v1 Announce Type: new Abstract: The majority of work on summarization evaluation focuses on general summary quality (e.g., ROUGE, BERTScore) or specific desired properties (e.g., reada
Its interesting to see the experiments on how to give AI a computer: Codex & Claude Code use your local machine, ChatGPT Work on the web (as well as Claude and ChatGPT) give the AI a one-time machine
arXiv:2608.14425v1 Announce Type: new Abstract: LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise. We introduce optstop, a
arXiv:2608.14150v1 Announce Type: new Abstract: The second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge evaluates two tasks over complete, unsegmented multilingual conversatio
arXiv:2608.14281v1 Announce Type: new Abstract: Forecasting crop growth across agricultural landscapes is important for improving the productivity, resilience, and operational management of farming sy
arXiv:2503.01236v3 Announce Type: replace-cross Abstract: This paper addresses fixed-graph terrain-aware path refinement, in which a global planner is restricted to a predefined route space and may re
arXiv:2608.14397v1 Announce Type: new Abstract: Zahavy [2026] argues that Large Language Models, despite their capabilities in induction and deduction, cannot perform the abductive 'Jump' that produce
arXiv:2608.14543v1 Announce Type: new Abstract: High-resolution image restoration from degraded inputs is challenging because it must preserve global structural consistency while recovering fine-grain
Many people’s blindness to this problem — AI failing to follow instructions — is like their initially blasé attitude towards hallucinations, with the same fantasy that core, endemic problems would rap
arXiv:2608.14367v1 Announce Type: cross Abstract: The early detection of delayed cases in business processes is a critical capability for organizations. Predictive process monitoring (PPM) supports th
Now that OpenCode plan isn't really worth it anymore, I'm looking for other options. I'm trying to figure out the Ollama Cloud subscription, specifically the $20 tier, because that's my budget for mon
Ollama has the best performance for deepseek v4 flash on average. For local only, you can try qwen3.8 that is optimized: Apple Silicon: ollama run qwen3.8:27b-mlx NVIDIA: ollama run qwen3.8:27b Benchm
I run locally Ollama + Qwen (tried multiples versions) and I connected it with MCP into my local SQLite database. When I ask a question, 90% of times I see it thinking and just stop and never answer t
arXiv:2608.13923v1 Announce Type: new Abstract: Open-vocabulary 3D scene graphs provide compact semantic memory for language-guided navigation, but mapped objects are often exposed through a single fu
arXiv:2608.14441v1 Announce Type: new Abstract: Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions a
arXiv:2608.13969v1 Announce Type: new Abstract: Prompt tuning adapts CLIP-based vision-language models with few trainable parameters, yet its predictions remain sensitive to the spatial sampling impos
arXiv:2608.14330v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables genome-wide gene expression profiling while preserving tissue architecture, but its cost and limited scalability re
arXiv:2608.14076v1 Announce Type: cross Abstract: Transition-state (TS) structures define the energetic barriers and mechanistic pathways of elementary chemical reactions, yet their identification rem
arXiv:2606.15158v2 Announce Type: replace Abstract: Reference-guided generation (e.g., object compositing, customization) has progressed rapidly, yet current pipelines share a fundamental limitation:
arXiv:2608.13959v1 Announce Type: new Abstract: Function calling is what the recent accounting of constrained generation explicitly sets aside: it finds the decoder's contribution small for format con
arXiv:2608.14104v1 Announce Type: cross Abstract: The Shapes Constraint Language (SHACL) is a W3C recommendation to express syntactic constraints, called shapes, on RDF graphs. SHACL validators are us
arXiv:2606.16316v2 Announce Type: replace-cross Abstract: Retrieving external knowledge is crucial for real-world tasks but remains difficult when queries and relevant knowledge are linked by implicit
arXiv:2510.02916v2 Announce Type: replace-cross Abstract: We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio fr
arXiv:2608.14046v1 Announce Type: new Abstract: In this work, we propose a source-agnostic framework that dynamically refines a binary mask throughout the reverse diffusion process by computing the di
arXiv:2608.14509v1 Announce Type: new Abstract: Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. This conflates two operations with d
the crux of why AI safety in systems that are built around LLMs is never going to work: they just can’t be trusted to follow instructions. Claude Code w/ Fable 5 has strict instructions in my project
arXiv:2608.13756v1 Announce Type: new Abstract: Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable. We test that assumption: holding the checkpoint
arXiv:2608.13817v1 Announce Type: cross Abstract: Human speech production is constrained by physiology, giving rise to characteristic temporal structure on acoustic signals. We hypothesise that these
Upped ChatGPT / Codex Subscription users limits to 900K in Hermes Agent, enjoy! GPT-5.6 Sol 1M in Codex. This used to only work for API keys, but we just flipped the switch and works for usage through
arXiv:2608.13791v1 Announce Type: cross Abstract: Positron emission tomography (PET) imaging suffers from limited spatial resolution and low signal-to-noise ratio, which can compromise quantitative ac
Watching a team of bots split up a large game project based entirely on their own experience and specialties with little input from me is pretty crazy (and just a bit creepy) to watch. I made a short
I don't understand this. Deepseek used half as many requests as GLM 5.2 for 3 times the usage, and I'm getting 10 times the requests from flash for like 1/100th the usage. I know Pro is supposed to be
I've pulled two models, `gemma4:latest` and `qwen3.5:latest`, and tried to use them in Android Studio and Itellij. I've followed the docs for both IDEs and they show in their chat features. My problem
llama : check LoRA tensor data is within file bounds (#27056) llama : check LoRA tensor data is within file bounds Update src/llama-adapter.cpp Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggi
chat: refactor handling supports_string_content / supports_typed_content (#27130) better supports_string_content cap detect test: add 'skip' messages_inp_normalizer Website: https://llama.app macOS/iO
Ollama sucks in terms of modaltiy options , model variety and is better in terms of usage limits . Is there any other alternative that could provide the same usage limits with better modality options
Ideas are the new bottleneck @akshaynathan_, Core Product Engineering, @OpenAI, interviewed by @swyx and @Vibhu (@LatentSpacepod) Summary: Akshay Nathan runs the productivity pillar at OpenAI, the tea
My local inference build, with: ASUS WS C422 PRO/SE 10-core Xeon W-2255 64GB ECC RAM 64GB VRAM Pimped case with TurboLEDz indicating the frequencies of the 10 xeon cores. Running llama.cpp with SYCL b
mtmd, common: various fixes (#27071) apply fixes cont revert gguf fix Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Int
common: migrate the deprecated --mmap/--no-mmap to --load-mode (#26934) Replace the deprecated --mmap, --no-mmap, --mlock, and --direct-io flags with the unified --load-mode argument across scripts, e
vulkan: add SHMEM_STRIDE_PAD/APPLY_SLM_A_RESHAPE for coopmat1 on Intel Xe (#25380) vulkan: add SHMEM_STRIDE_PAD/APPLY_SLM_A_RESHAPE for coopmat mul_mm on Intel Xe vulkan: fix shmem estimate for Intel
fix: check gguf array type before reading (#27075) fix: check gguf array type before reading update skill Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64,
vendor: update BoringSSL to 0.20260813.0 (#27099) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFrame