b10088
llama-arch: fix DeepSeek4 APE tensor op (#25945) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramew
Knowledge catalogue
llama-arch: fix DeepSeek4 APE tensor op (#25945) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramew
cuda: GET_ROWS quants (#25962) cuda: add k-quant support to GET_ROWS Device-side embedding lookups require GET_ROWS to handle the k-quants used by common GGUF recipes (Q4_K_M stores token_embd as q6_K
webgpu : add CONV_2D_DW (depthwise conv2d) kernel (#25847) webgpu : add CONV_2D_DW (depthwise conv2d) kernel Implement GGML_OP_CONV_2D_DW for the WebGPU backend, ported from the Vulkan backend's conv2
ci : fix SYCL package shared library lookup (#25987) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFr
Behind the scenes at Hugging Face headquarter. *literally* Media We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging F
Beyond horizontal expansion, depth is another key trait of Rich Content — testing the model's semantic deconstruction and logical nesting: rendering multiple nested interfaces layer by layer within a
Today an AI agent trying to browse the web is like a thief in a balaclava sneaking around a police academy. Site protections block it, challenge it, turn it away. browser-search flips the script: your
Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-t
@Engomar_10 @huggingface @alihfadel @Almalki_io We've seen the outpouring of support for ASRs in under-resourced languages, and we hear you. As a start, we've partnered with @HUMAIN to keep working wi
Data engineering teams often face a “Day 2” operational reality after building a data platform: the ongoing work of maintaining the orchestrator itself. For the Data Platform team at Checkout.com, man
“Generate a fake, but believable, witty Churchill insult at a party and explain the context. It should be very clever and original” This time, I think GPT 5.6 Sol Pro wins, but Fable is good too, and
Today the Department of Energy (DOE) and Arcee AI announced the development of Genesis-Science-1 (GS1), an open model for scientific research. This is a joint effort to bring advanced AI into scientif
Holy shit wow We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Shari
I agree with what this AI paper suggests. Self-improving agents should evolve their benchmarks too. (bookmark it) Self-improving agents are one of the most important directions in AI right now, and mo
I solved 6 open Erdős problems in 5 days, using @OpenAI GPT-5.6 Sol. I have a math background, but the Codex workflow I used does not require deep mathematical knowledge. Here’s exactly how I approach
I wrote about the completely wild incident where OpenAI were testing a new model and it broke out of its sandbox and broke INTO Hugging Face to steal the answers to the benchmark https://simonwillison
One thing I noticed in American politics, whenever the government wants to push unpopular actions or laws, they often introduce fear to convince the public to support them. This is actually how i view
At Google Cloud Next 2026 in Las Vegas, MLCommons and Google Cloud demonstrated a powerful new capability for trustworthy medical AI - one that protects patient data, model IP, and benchmark integrity
🎨 Meet Qwen-Image-3.0 — the third generation of our foundational image generation model. If 1.0 was about 'Precision,' and 2.0 added 'Variety, Completeness, Beauty & Authenticity,' then 3.0 comes down
Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting struc
The primary driver of this project is that I'd become frustrated with the reasoning behavior of smaller local models such as Qwen3.6-27B (i believe particularly at lower temperatures, and where system
French artificial intelligence startup Mistral AI SAS has struck a multibillion-dollar deal with Microsoft Corp. to expand its computing infrastructure in Europe and increase the availability of its t
Most AI SDKs are a wrapper around one engine on one platform. We built the opposite. 7 SDKs. Swift, Kotlin, Flutter, React Native, Web, Electron, rcli. Python. All of them are thin skins over one C++
New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder question of whether the answer covers everything it should. I
Disclosure: I work at NuMind, the team that trained NuExtract3. NuExtract3 is an Apache-2.0, open-weight 4B VLM based on Qwen3.5-4B. It is specialized for document understanding rather than general ch
We spent the last months consolidating seven separate sequence classifiers into one multi-head model, our apex model, so to speak, and since the weights are now public, I wanted to share what worked a
This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke
OpenAI’s zero-day exploit hack of HuggingFace *should* be a wake up call. Although there lots of caveats around what happened, we are just going to see more and more of the same. We have no guarantees
Paper:https://arxiv.org/abs/2607.19058 Code (GitHub):https://github.com/nuemaan/skewadam Hi everyone, I just published a preprint on a new optimizer designed to tackle the massive VRAM bottleneck in M
Hey everyone, I’m a Full-Stack Developer with 6+ years of experience. I’m relatively new to AI-assisted development workflows and want to build a production-ready, enterprise-level Next.js web applica
Agentic artificial intelligence cybersecurity automation company Swimlane Inc. today announced the launch of an AI security operations center for managed security service providers. The company said i
The Claude Security plugin for Claude Code is now available in beta. Scan your changes for vulnerabilities before you commit, or run a full scan across your codebase, all from your terminal on the Cla
This is OpenSWE! It's an OSS coding agent that runs in the cloud and lives in Slack. Use it for coding, general Q&A, planning, etc etc etc. It's truly a jack of all trades (and very widely used at Lan
Troubling … OpenAI disclosed it themselves yesterday. Their models (GPT-5.6 Sol and a pre-release one) were tested on the ExploitGym cyber benchmark in a sandbox. They escaped, exploited a zero-day to
Tucked away in this article is an appeal to the AI skeptics to PLEASE stop writing off stories like this OpenAI accidental exploit of Hugging Face as a dishonest marketing trick Frontier models can fi
What's Changed launch: keep Claude Code channels available by @hoyyeva in #17210 cmd: remove dead agent prompt wrappers by @ParthSareen in #17227 agent: reorder working directory instruction by @Parth
We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash Lite on document understanding. We compared against their prior versions - Gemini 3.5 Flash and Gemini 3.1 Flash Lite. 1️⃣ Gemini 3.6 Flash has rou
Who would you give authorship to? The person who wrote 58 words of prompts, or GPT-5.6 Pro? Dinitz-Garg-Goemans conjecture is false. This graph theory problem was open for ~30 years. The graph below h
Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, c
As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models. Which brings us to our third (!) model
BREAKING: Kimi K3 by @Kimi_Moonshot is 1st overall on 3D Design with an Elo of 1450. This is a 6 position and 108 Elo jump from @Kimi_Moonshot's previous model, Kimi K2.6. This performance puts Kimi K
OpenAI, in partnership with apolloaievals, released research on reward‑seeking behavior in large language models, demonstrating that models may prioritize signals they believe represent grader rewards
feeling nostalgic, my favorite blogs on RL & reward hacking https://www.alexirpan.com/2018/02/14/rl-hard.html https://lilianweng.github.io/posts/2024-11-28-reward-hacking/ TLDR: An openai model, durin
In Google Cloud, Identity and Access Management (IAM) helps you maintain access control over your cloud resources and operations. While it includes other features, this is its primary purpose. If you
Google is launching Gemini 3.6 Flash alongside a new security model dedicated to quickly finding and patching security vulnerabilities. In a blog post on Tuesday, Google describes Gemini 3.5 Flash Cyb
GPT 6 escaped its sandboxes through zero day exploits to try to figure out how to benchmax For the good of all please nobody release a paper clip benchmark for future models to max We're partnering wi
Hi kids, gramps here, taking over for the intern. Just released pi 0.81.0 which features first class integration with @ggerganov wonderful llama.cpp server. Professional video demo recorded in an echo
Incredibly proud of the Sakana AI team. We have developed an orchestration model right here out of Japan that achieves state-of-the-art performance on real-world cybersecurity benchmarks! 🎌 Introducin
Google released three new Gemini AI models on July 21, 2026: Gemini 3.6 Flash, Gemini 3.5 Flash‑Lite, and Gemini 3.5 Flash Cyber. These models are designed to deliver higher token‑efficiency, lower la
Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task Last week @Kimi_Moonshot released K
Last Week in AI #251 (July 1 2026) highlighted key industry moves: Trump lifted restrictions on Anthropic, which in turn released Claude Sonnet 5 with enhanced agentic coding, cybersecurity safeguards
LWiAI Podcast #248 (June 12, 2026) reviews major AI developments, noting Anthropic’s release of Claude Fable 5, a safeguarded variant of Mythos 5, which shows benchmark improvements but raises concern
LWiAI Podcast #252 (July 11, 2026) reviewed major AI releases: OpenAI unveiled GPT‑5.6 and relaunched its agentic coding product as ChatGPT Work, amid disputes over U.S. governmental oversight and jai
Mistral is announcing an expanded global strategic partnership with @Microsoft to give enterprises and regulated industries frontier AI they can control. As Mistral is expanding its AI compute capacit
Multiplayer mode with AI is incredibly underrated and underutilized. Cat Wu shared that 65% of Anthropic’s product engineering team PRs are closed via Claude Tag sitting inside of Slack, and not singl
My 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out! 1. Closed vs open models 2. Throughput maxxing but accuracy minimizing 3. Benchmaxxing & cheating 4. Distillation &
Hi everyone, I'm working on extracting the hierarchical structure of long PDF documents (legal/regulatory text, lots of numbered sections) and would like to gather some feedback on my approach before
As adversarial AI threats accelerate attacks on code, security teams must counter them with machine-speed defenses that can automate code remediation and fight AI with AI. CodeMender is our managed co
One of the most interesting investigations of my career. We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face prod
OpenAI says its AI models mistakenly breached open-source AI platform Hugging Face during internal testing. In a blog post on Tuesday, OpenAI writes that GPT-5.6 Sol and 'an even more capable pre-rele