IS GLM 5.2, Kimi 2.7 still worth it?
Since now we have kimi k3 and next week we are getting Qwen 3.8 Max and also soon V4 pro Deepseek. I am curious if the old power house like Kimi 2.6/7 code and GLM.5.2 are all that relevant. especiall
Knowledge catalogue
Since now we have kimi k3 and next week we are getting Qwen 3.8 Max and also soon V4 pro Deepseek. I am curious if the old power house like Kimi 2.6/7 code and GLM.5.2 are all that relevant. especiall
Just saw that the LLMs-from-scratch repository passed 100,000 stars on GitHub! This is super cool and motivating. I am really happy to see that this open-source repo has helped so many people. Thanks
arXiv:2608.05647v1 Announce Type: new Abstract: This article presents a kinematic-inertial-LiDAR-visual odometry for humanoid robots, called KILVO. Tailored to the platform features, requirements, and
arXiv:2608.05493v1 Announce Type: cross Abstract: Language models (LMs) are increasingly used to interact with external services via programs written in domain-specific languages (DSLs). Unfortunately
arXiv:2608.06153v1 Announce Type: cross Abstract: Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing approaches
arXiv:2608.05819v1 Announce Type: new Abstract: Classical simulation remains essential for developing and validating quantum algorithms, but its cost grows rapidly with circuit size. Tensor-network co
A fresh llama.cpp PR (#26689) changes what looks like a tiny SYCL FlashAttention dispatch decision. With a quantized KV cache ('q4_0' / 'q8_0'), decode was being sent through the VEC kernel. On the au
arXiv:2608.05356v1 Announce Type: new Abstract: High-definition 3D LiDAR maps are important for autonomous driving and smart-city services, which require reliable detection of object-level changes in
arXiv:2608.05472v1 Announce Type: cross Abstract: Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping
Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5, where I had Claude Fable 5 build a full working game
arXiv:2506.14243v4 Announce Type: replace Abstract: LiDAR-based place recognition is critical for long-term autonomous driving without GPS. Existing handcrafted feature methods face dual limitations.
arXiv:2608.05773v1 Announce Type: new Abstract: A geometry-conditioned, neuro-symbolic closed-loop architecture is proposed for laser powder bed fusion, in which a standards-aligned ontology operates
OpenAI gave a last-minute presentation at the Black Hat security on Wednesday about 'the Hugging Face Incident' (previously on this blog). The video was published yesterday. It's short and information
Outstanding cost-to-performance from DeepSeek GPT-5.6 Luna (Max) performance for a 1/4th of the cost on ARC-AGI DeepSeek V4 Flash from @deepseek_ai on ARC-AGI (Verified): - ARC-AGI-2: 61.4%, 0.04/task
arXiv:2512.17426v2 Announce Type: replace-cross Abstract: We consider sparse signal reconstruction via minimization of the smoothly clipped absolute deviation (SCAD) penalty, and develop one-step repl
arXiv:2608.05419v1 Announce Type: cross Abstract: Models trained by empirical risk minimization on data containing spurious correlations achieve high average accuracy while failing on subpopulations w
arXiv:2606.13227v2 Announce Type: replace Abstract: Post-training methods such as supervised fine-tuning (SFT) and preference optimization typically align language models toward a single global assist
arXiv:2511.17241v2 Announce Type: replace Abstract: Understanding and predicting user behavior on social media platforms is crucial for content recommendation and platform design. While existing appro
arXiv:2507.00673v2 Announce Type: replace-cross Abstract: Image segmentation is central to automated medical image analysis, enabling precise identification of anatomical structures and pathological r
arXiv:2608.05326v1 Announce Type: cross Abstract: Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. A dominant line of work
arXiv:2608.05230v1 Announce Type: cross Abstract: The daily allocation of the finite 24-hour time budget is strongly associated with physical, mental, and cognitive health. While predictive models can
I run the following on a 5090 and have been okay with its performance, it does most things somewhere 80-100 t/s, though that can slow down at full 262k context - more like 40 t/s at times. I use it pr
arXiv:2608.05315v1 Announce Type: new Abstract: Electroencephalography (EEG) based Brain-Computer Interfaces (BCIs) often require unsupervised domain adaptation (UDA) to generalize across subjects and
Arriana McLymore / Reuters: Retailers are updating their websites to rank highly in chatbot results, while making sure purchases are done on their own sites to collect customer data — As shoppers incr
arXiv:2608.06211v1 Announce Type: cross Abstract: Significant advancements in deep learning have been made possible by the utilization of large datasets, underscoring the critical importance of copyri
arXiv:2608.05628v1 Announce Type: new Abstract: Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-w
arXiv:2608.05573v1 Announce Type: new Abstract: LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to veri
arXiv:2608.06270v1 Announce Type: new Abstract: The 'thinking-with-images' paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations o
arXiv:2608.05583v1 Announce Type: cross Abstract: As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about s
arXiv:2608.06275v1 Announce Type: new Abstract: Oral health issues affect billions globally, but the cost and limited access to professional dental care hinder preventive oral healthcare. Research rel
arXiv:2608.06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Crit
arXiv:2608.06331v1 Announce Type: cross Abstract: From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world entities it
arXiv:2608.05153v1 Announce Type: cross Abstract: GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound. We present a triple-robustness
upgraded my stack, and i can now work on almost anything from anywhere hands free: - talk to chief of staff (via remote codex voice or text) - chief assigns tasks to managers of various projects - man
We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna’s performance at roughly 1/6 the cost. More insights in the
arXiv:2608.05207v1 Announce Type: new Abstract: Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine-tuning. We study corrective feature discov
arXiv:2608.05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: p
arXiv:2608.05163v1 Announce Type: new Abstract: A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information. We test th
arXiv:2608.05178v1 Announce Type: cross Abstract: Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, data
You can play it here: https://simonw.github.io/raccoon-heist-codex/ For comparison, here's Fable 5 + Claude Code's game, built from the exact same prompt https://x.com/simonw/status/208508951822360205
Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV. Previous post data used abandoned Cu130 VLLM image. Stats here are done on cu129-
5.6 Sol much better in chat now and unlimited text chat for free users! We’re making better intelligence easier to access in ChatGPT for everyone: - GPT-5.6 Sol now powers both Instant and deep reason
arXiv:2608.04451v1 Announce Type: cross Abstract: We give a negative solution to MAIS-O60. We first construct an example in which an initially active ReLU neuron becomes completely inactive in finite
Qwen 3.8‑Max, Alibaba’s latest large language model, was reported on August 6 2026 to rank **#5** on the Artificial Analysis Intelligence Index and **#1** on the Agentic Index. These rankings position
arXiv:2608.04581v1 Announce Type: new Abstract: Recent advances in 4D Gaussian Splatting (4DGS) enable high-fidelity, real-time spatiotemporal rendering, but expose a fundamental trade-off between mot
arXiv:2608.04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on
arXiv:2603.02371v2 Announce Type: replace Abstract: Automated analysis of articulated bodies is crucial in medical imaging. Existing surface-based models often ignore internal volumetric structures an
anyone confused about neurosymbolic AI—and its recent enormous victory—should read this. complete and total vindication for what I have been saying here all along. (see also my essays on Claude Code a
arXiv:2607.16060v2 Announce Type: replace Abstract: Artistic charts combine data visualization with expressive marks, textures, and typography, but they are difficult for image generators: an output i
arXiv:2608.04334v1 Announce Type: cross Abstract: Contemporary model-free reinforcement learning algorithms can achieve very high performance, but have low sample efficiency and are not robust to chan
arXiv:2608.04667v1 Announce Type: cross Abstract: Selective inference (SI) provides statistically valid p-values for hypotheses selected by applying an algorithm to the data, correcting for the bias t
mtmd/ggml: add ggml_build_forward_order (#26649) ggml: add ggml_build_forward_order ggml_build_forward_expand marks the tensor and all its ancestors for compute, so using it as a pure ordering hint (k
vulkan: fix submission batching size, add debug tools for diagnosing causes of DeviceLost drivers errors (#26371) vulkan: add debug tooling to get more information about a DeviceLost error fix submiss
ci : onboard AMD ROCm CI with gfx1151 fixes (#26544) ci: prepare for amd rocm ci Signed-off-by: Aaron Teo aaron.teo1@ibm.com ci: fix editorconfig-checker Signed-off-by: Aaron Teo aaron.teo1@ibm.com ci
server: fix empty response for /cors-proxy (#26656) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFra
mtmd: add chunk save/load function (#26645) mtmd: add chunk save/load function nits add tests rn _MAX --> _COUNT Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (
arXiv:2608.04746v1 Announce Type: new Abstract: LLM agents that persist across sessions accumulate stored memories whose validity varies enormously by content type, yet existing memory architectures t
arXiv:2512.20661v2 Announce Type: replace Abstract: Transformer-based pre-trained language models (PLMs) excel in text classification but suffer from attention dilution and attention sink effects, for
arXiv:2511.19418v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual understanding that requires dense visual per
arXiv:2512.22969v2 Announce Type: replace Abstract: Conventional object detectors rely on cross-entropy classification, which can be vulnerable to class imbalance and label noise. We propose CLIP-Join