Auditable Agents
arXiv:2604.05485v2 Announce Type: replace Abstract: LLM agents call tools, query databases, delegate tasks, and trigger external side effects. Once an agent system can act in the world, the question i
Knowledge catalogue
arXiv:2604.05485v2 Announce Type: replace Abstract: LLM agents call tools, query databases, delegate tasks, and trigger external side effects. Once an agent system can act in the world, the question i
dflash : clarify output logging of target_layer_ids (#27013) This commit tries to make the logging of target_layer_ids a bit clearer and easier to read. Currently the output generated looks like this:
sycl: fuse the gated-delta-net state writeback cpy (#26643) Port of #23940. Arc Pro B70, Qwen 3.6 27B Q4_K - Medium (48 of its 64 blocks run gated_delta_net), -ngl 99 -fa 1 -ctk f16 -ctv f16 -b 2048 -
ggml: force single thread on wasi (#25686) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Li
sycl: fuse mul_mat(gate) + mul_mat(up) + GLU for q4_K dense FFN (#26779) Measured on Arc Pro B70 (Battlemage, Level Zero), llama-bench -r 20, two interleaved rounds, tg128: qwen2.5-3B-Instruct Q4_K_M
tests : replace personal home directory paths with generic placeholders (#27043) Scrub developer-specific /home// paths from example docs and test fixtures so they don't leak into the tree. examples/t
server: allow accessing /metrics and /slots during llama_decode() (#27041) server_queue::worker call llama_decode inside yield_to_queue also handle process_mtmd_chunk clean up nits rm test Website: ht
llama : allow virtual igpu devices (#26953) llama : allow virtual igpu devices cont : better comment Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, Kleid
ggml : recurrent state rollback for ggml_ssm_scan (#26623) Initial changes for Recurrent state rollback for nemotron for cpu and cuda Removing CPU RS rollback. Will enable it in subsequent PRs additio
sync : ggml Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu ar
jinja : fix quadratic cost in gather_string_parts (#27034) jinja : fix quadratic cost in gather_string_parts fix some comments remove test Website: https://llama.app macOS/iOS: macOS Apple Silicon (ar
arXiv:2608.13524v1 Announce Type: new Abstract: Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters furt
arXiv:2608.12753v1 Announce Type: new Abstract: We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmet
arXiv:2608.13474v1 Announce Type: new Abstract: Vision-language-action models (VLAs) are moving rapidly towards deployment as general-purpose manipulation policies, but we currently lack basic tools f
arXiv:2510.24232v3 Announce Type: replace Abstract: To improve detection robustness in adverse conditions (e.g., haze and low light), image restoration is commonly applied as a pre-processing step to
arXiv:2608.12921v1 Announce Type: cross Abstract: The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology
arXiv:2608.12737v1 Announce Type: new Abstract: Deep representation learning has primarily focused on how features evolve across network layers, while largely overlooking the structured geometry embed
arXiv:2608.12663v1 Announce Type: cross Abstract: Wildfire susceptibility mapping typically relies on physical variables assembled from multiple remote-sensing, climate, and geospatial products. Alpha
arXiv:2510.03880v2 Announce Type: replace Abstract: In recent years, with the rapid development of large multimodal models (LMMs), explainable image quality assessment (IQA) has attracted increasing a
while having much fun testing LLMs Houdini-like attitudes and abilities to evade, excalate and escape from carefully reciprocally arranged security enhancing sandboxing VMs, containers, namespaces and
arXiv:2608.13096v1 Announce Type: new Abstract: Limit order book (LOB) simulators are most useful to practitioners when they combine realistic market dynamics, computationally efficient sampling, cont
arXiv:2608.13229v1 Announce Type: cross Abstract: We present the mathematical foundations of linear independent component analysis (ICA) models based on standard literature in a self-contained note. I
arXiv:2608.13018v1 Announce Type: new Abstract: Standard probabilistic logic programming frameworks typically rely on grounding logic programs into discrete propositional representations. This operati
From idea to implementation in one go.🏃♀️ Max-level intelligence, served fresh on Day 0. Qwen3.8-2.4T-A95B is live on SiliconFlow. Thanks! @SiliconFlowAI 🚀 Day-0 Support! @Alibaba_Qwen has open-sourc
arXiv:2602.23672v2 Announce Type: replace-cross Abstract: This study proposes a General Bayes framework for policy learning. We consider decision problems in which a decision-maker chooses an action f
Emma Roth / The Verge: Google adds a toggle in Gemini and Flow to remove visible watermarks from AI-generated images, videos, and music; SynthID watermarks and C2PA metadata remain — Google will still
arXiv:2510.05678v2 Announce Type: replace-cross Abstract: While large language models (LLMs) have achieved notable progress in multilingual settings, their performance remains uneven across languages
Grok 4.6 ranks #1 on CursorBench for real-world coding Grok 4.6 just ranked #1 on CursorBench 3.2 Outperforming Claude Fable 5, Opus 5 and GPT-5.6 Sol on real-world coding performance And what makes t
arXiv:2608.12335v1 Announce Type: new Abstract: Financial question answering over annual reports requires more than retrieving semantically similar passages. It often involves identifying relevant com
arXiv:2608.13215v1 Announce Type: new Abstract: Forecasting the long-horizon evolution of mechanical systems from position-only observations is a pivotal yet difficult task, as hidden velocities and t
arXiv:2608.13555v1 Announce Type: cross Abstract: Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kin
arXiv:2608.12915v1 Announce Type: cross Abstract: The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions sha
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting
arXiv:2608.13296v1 Announce Type: cross Abstract: Existing global optimization benchmark suites are of a moderate size and are based on a small number of analytical functions that date back even to th
arXiv:2608.13450v1 Announce Type: cross Abstract: Autonomous vehicles depend on large safety-critical software stacks, where weaknesses reachable from adversarial inputs may affect steering, braking,
arXiv:2608.12626v1 Announce Type: cross Abstract: Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals. In these settings, fin
arXiv:2412.17228v4 Announce Type: replace Abstract: Background: Clinical trials are essential to advancing cancer treatments, but fewer than 10% of adults with cancer enroll in therapeutic trials. Ope
arXiv:2603.19703v2 Announce Type: replace-cross Abstract: Estimating covariance matrices is fundamental to a wide range of statistical applications. This paper studies minimax and adaptive estimation
arXiv:2608.12371v1 Announce Type: new Abstract: Stream-processing systems increasingly operate across heterogeneous mobile edge--cloud infrastructures, where workload volatility, resource contention,
My Al Agent Workshop with @mcuban might be the best workshop we've ever had. Over 20,000 people registered and my inbox is filled with people making REAL CHANGE in the way they work and what they're a
new langchain oss release! What's new in LangChain? 🚀 🔌 Support for OpenAI's 3.0 SDK (using httpx2) ✨ Support for gemini-3.7-flash Plus a wave of core reliability fixes from external contributors: 🛠️
arXiv:2608.12582v1 Announce Type: cross Abstract: AI journaling tools can tailor prompts to a person's own sensed behavior, but it is unclear which behaviors respond to them. We analyzed 369 journal e
arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today. Because of their ubiquity and computa
arXiv:2608.13510v1 Announce Type: cross Abstract: Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. However, their achievable performance
OpenAI and Anthropic have slashed prices for their flagship models—OpenAI cut GPT‑5.6 Luna by 80 % and Anthropic launched Claude Opus 5 at half the cost of Fable 5—in an effort to stem customer defect
arXiv:2502.08004v2 Announce Type: replace-cross Abstract: Simulation-based inference (SBI) is a method to perform inference on a variety of complex scientific models with challenging inference (invers
arXiv:2608.12789v1 Announce Type: cross Abstract: Tool-using agents consume external data from sources with different levels of trust, yet tool responses rarely identify who produced each component or
arXiv:2608.11506v1 Announce Type: cross Abstract: Adaptive behavior under partial observability depends on internal organization that carries information beyond the current observation. Drawing on Bar
I've downloaded qwen3.8:27b-mlx through ollama and attempting to run it with Claude Code. On any command, I'm getting 'API Error: 500 system message must be at the beginning. This is a server-side iss
arXiv:2608.12905v1 Announce Type: new Abstract: The behavior of contemporary generative Large Language Models (LLMs) is directly shaped by prompts, unstructured texts that describe the desired output
So, I have been dreaming of getting 17 tokens per second using my RTX 3050 6GB version on a decent context window for Hermes needed above 60k. The hope is that has was a 22GB of DDR 4, hoping they can
Qwen 3.8-Max is live on Modal, with the full 1M context window and a custom DFlash speculator under the hood. Love seeing our launch partner Modal go all in from Day 0!⚡️ 2.4T parameters. 1M context.
Qwen Live now, an appetizer 👀 Ready for Qwen Live EP2?We will start at 10:00AM! UTC+8 https://x.com/i/broadcasts/1mGPaZanWBqJN Qwen Cloud:https://www.qwencloud.com/?utm_content=g_20000002312 #QwenClou
Qwen3.8-Max is live on Together AI. Together AI is with us as a Day 0 launch partner, and we couldn’t ask for a better name to share Day 0 with. 2.4T parameters, 95B active, 1M context — all together
Regular reminder -- the set of public ARC 3 games is called 'demonstration set', not 'eval set' nor 'training set'. It is not meant to be used as training data, and it is not meant to be used as an ev
arXiv:2608.12879v1 Announce Type: new Abstract: Fractional partial differential equations describe nonlocal dynamics, but discovering them from noisy data is difficult because fractional differentiati
arXiv:2608.13108v1 Announce Type: new Abstract: Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistenc
arXiv:2608.13538v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features stil
arXiv:2608.13541v1 Announce Type: new Abstract: Part-aware 3D generation aims to create digital assets that are coherent as complete objects while exposing structural parts for editing, material assig
arXiv:2608.13368v1 Announce Type: cross Abstract: This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Netwo