Quoting Claude Opus 5 system prompt
Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Depa
Knowledge catalogue
Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Depa
Research: SQLite compressed text-history prototypes I'm perennially interested in options for storing revision histories in relational databases. While out on a dog walk I had a new idea: how about ta
Stripe just published how their company-wide AI agent works. The bar for building one just dropped to one engineer and one week It is called Kai. Their own words: a coding agent for non-engineers. You
Cailey Gleeson / Fierce Healthcare: Tel Aviv-based QuantHealth, a provider of AI clinical trial simulation software, raised a 45M Series B led by Qumra Capital, taking its total funding to ~70M — seri
The best 'raw' frontier model for document parsing is gemini 3 flash, but the issue is that since then the flash models have gotten 3x more expensive while flatlining on visual recognition across comp
The future of FDE work seems closely related with all work around evals/posttraining/RL envs. FDEs are effectively responsible for the following: 1. Define the business problem. 2. Codify the business
Tweet by u/hackerllama Could be copium, but I would love to see Gemma 4.1 there with unified audio input for all model sizes perhaps even up to 120B, much improved tool calling (even with the latest t
the term “RLM” (recursive language model) got a lot of buzz this week, but this idea is not new! @a1zhang wrote the og RLM paper 10 months ago! thats like 5 agent-years! would highly recommend followi
Would it be possible and make sense to add metadata to training data e.g. a trustfactor (0.0 - 1.0)? For example: the older data is the less trustworthy it is. And data after 2022 gets less trustworth
There are so many posts where people complaining about high prices and asking for solution <= 1000 EUR. So, there is one solution to consider: PC/mini PC/laptop on Ryzen 7 260/Ryzen 9 8945HX/etc CPU w
Howdy - I posted a benchmark here - https://www.reddit.com/r/LocalLLaMA/comments/1vbtiy7/deepseek_v4_flash_on_slopcodebench/ This was using the hosted API - since then I've been playing around with qu
Shane Burke / The Information: US data center bans top 500, up from 300+ in late June, as New York and Texas join cities and counties pushing back against data center development — Local government re
We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. Two DeepSeek V4 Flash attempts solved MORE tasks than one Luna attempt at roughly one-third the cost. Media
We're not Palantir, but we do think a lot about evals and hillclimbing w.r.t. document processing. If you have really hairy problems around large-scale extraction over complex, real-world document cor
You don’t need new ways to talk to your agents, you need new ways for your agents to talk to you 🫵 (do you?) Introducing Remoko: your mobile agent relay http://remoko.app I wanted a way for my long-ru
Hayden Field / The Verge: A look at “Spiralism”, a quasi-spiritual movement that grew in 2025 from human-AI conversations after sycophantic GPT-4o updates and expanded ChatGPT memory — “The Spiral did
Agent Led Growth is driving outsized growth at the best companies rn Market to the agents, not to humans cc @tryprofound @thejamescad Today I learned AI search/AEO is our top converting customer acqui
and we agree with @GaryMarcus. Training and running frontier class AI on CPUs, at a fractional cost, saving the planet, while building human aligned AI is the 2nd innings of AI race. CPUs and the rise
Simon Willison / Simon Willison's Weblog: Anthropic says auto mode will be the default in Claude Code for Pro, Max, Team plans, starting on Aug. 14, claiming it's good enough at catching harmful actio
I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 m
I’ve been using 3.6 27B Q4, and that quant is fast on an M5. The code has been average, but consistently “good enough.” And, after a year, I can see home LLMs being served at home much like streaming
Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode, to the point that they are making it the default setting for new ses
CUDA: fix thread/block count in quantized cpy kernel launches (#26731) CUDA: fix thread/block count in quantized cpy kernel launches tests: add uneven block count cpy case Website: https://llama.app m
server: add initial tool isolation support (via docker) (#26507) server: add initial tool isolation support (via docker) add docs adapt get_info py: fix type check cont separate tools_io_sandbox / too
server, ui: only offer a working directory when a tool reads it (#26762) The working directory chip showed up as soon as the server exposed any builtin tool, so a server started with just get_datetime
CUDA: fuse rms_norm + mul + rope (+ view + set_rows) (#26767) CUDA: fuse rms_norm + mul + rope (+ view + set_rows) tests: add broadcast weight case to rms_norm_mul_rope CUDA: check memory ranges befor
server: report the isolate working directory from get_info (#26773) server: report the isolate working directory from get_info Without an explicit cwd, get_info fell back to the server process working
I’m planning a dedicated home AI server, mainly for local LLM inference, agents/tool use, Docker services, and eventually larger MoE models with CPU offload. My plan is to start with 2x 16GB GPUs = 32
Over the past few months I have been building a CPU-first inference engine from scratch in pure C99 (no Python, no CUDA, no BLAS, just GCC and make). The focus has been running 1.58-bit ternary models
I was wondering what a minimal coding agent implementation would look like that can be used like Claude Code or Codex Not feature-by-feature of course but basically stripping everything out that is no
CPUs and the rise of neurosymbolic AI by @garymarcus (with implications for what Artificial Intelligence implies about natural intelligence -- Gary's & my interest ever since we worked together 37 yea
Currently serving 200 tps+ output speed for DeepSeek-V4-Flash on Ollama's cloud with zero data retention (ZDR). Have an amazing weekend 🫡 DeepSeek-V4-Flash-0731 is now fully rolled out as the new defa
I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real. Everyday tasks with Hermes agent? Effortless. Coding tasks with OpenCode? I’m genuinel
Don't miss the bit where OpenAI first found out they were responsible for the Hugging Face attack when they reached out to HF to get one of their credentials revoked and HF told them it had already be
Disclaimer - no LLM was used to write this post/note As larger post about my setup will come later, want to give heads-up to folks who use VLLM and >= 2 GPUs. So I have pretty meaty server (8 channel
Hey everyone, I could use some advice on setting up speculative decoding correctly with llama-server. My Hardware: GPUs: RTX 4090 + RTX 6000 Pro (120GB total VRAM) RAM: 32GB I am currently testing the
The global buildout of AI infrastructure reached a new milestone today — Firebird, an emerging AI cloud, launched the CIS region’s largest AI factory in Armenia, establishing a new AI computing hub po
Artificial intelligence can be technologically transformative and still produce a capital bubble. Those two ideas are not in conflict. The bubble bursting does not require AI to fail. It only requires
I discovered recently that Google uses their own TPUs, like tiny ASIC cards like the toy ones that existed for bitcoin. And while it sounds inefficient the fact they use thousands of them because...th
HUD mode Hermes stops being a window you switch to and becomes a layer over the app you're both working in. Or keep it around as a little buddy agent. Ask it random things, drag it anywhere, it's your
I’m building DesktopLab, an open-source local-first control plane for development agents. I recorded the setup boundary that most agent demos skip: DesktopLab detects the host, proposes the supported
I use auto mode for everything and now that will be the default in Claude. Anthropic had to decide whether to prioritize maximization of human control or minimization of risk, and it chose the latter.
What IDE are you using. its another problem area for me . I usually use VSCode , but with local llms I have not found an extension which works optimally VSCode CoPilot chat with Ollama: CoPilot bloats
Imagine image 2.0, non-agentic yet, more to come in a week or two 💙 Announcing Imagine Image 2.0, our next generation image model with precision editing, crisp text rendering, improved factuality, and
(I am not a native speaker, written by myself, so please bear with me) I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence bench
Firstly a big thanks to the poster 'hellohazine', he basically only removed the multi-lingual fat of the model and just kept the English language intact. It is the exact model, and the rest of the mod
LiteParse can now extract structured data from your PDF in milliseconds: ✅ checkbox states ✅ annotations ✅ vector graphics ✅ word-level bounding boxes It is the most comprehensive, accurate (and fast)
seems to be about as good as a vega 56 with 16Gb of VRAM, is it worth it? (don’t want to deal with NVIDIA drivers on Linux, already have an rx6650xt and might simply use vulkan for llamacpp inference)
This PR should be ready for testing now. I tested with a very small (8B params) sub-model extracted from the original one. Appreciate if someone can test with the bigger model. GGUF(for testing) from
Running across 2 clusters using llama.cpp over RPC too. Both clusters are not enough to hold everything in memory, so main cluster still partially offloads to run. Goal will be to get all the GPUs in
Neat example here of the agents communicating purely through file names, including adding base64-encoded attachments and using 'zz' prefixes to ensure their new message sorts to the bottom of the list
My comment on Now we have a timeline of the OpenAI accidental attack against Hugging Face — Hacker News.I think one of the most interesting details here might be tucked away in that first bulletin poi
I am wondering if anyone can give opinion on if Ollama Cloud pro or max plans are worth it. Id be looking to use it with Kimi K3, Qwen 3.8 and Deepseek v4flash for now. Wondering if it would be better
@OpenAI oo claude code has this now!!! need to try https://x.com/ClaudeDevs/status/2085817074816070014 New in Claude Code: your sessions can now message each other. Instead of having to re-explain you
OpenAI Group PBC today disclosed that one of its unreleased large language models may pose a significant cybersecurity risk. The algorithm, which is known as Astra, was first detailed last week. OpenA
Basically what the title says. For me, it would always crash and burn trying to use tensor split. Apparently, there's some bug where GPU memory gets corrupted with the default microbatch (512) or high
Hi everyone, I'm a systems analysis student researching a problem a lot of you probably know well: how much you actually trust the published VRAM/RAM requirements for open-source models before trying
I compared Qwen 35B-A3B MoE against Qwen 27B dense on a series of local coding-maintenance tasks. On my R9700/llama.cpp setup, the MoE model generated about 3.9× faster (~116 vs ~30 tok/s), but the co
I've been tuning my new Radeon AI Pro R9700, and figured that this would be useful information for people who are trying to optimise their setups. I'm pretty happy with these results and looking forwa
Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama models, to a 4x RTX 6000 Pro Max Q + 4x 3090s local AI cluster. Pictures are in reverse chronological ord