Consistent Diffusion Language Models
arXiv:2605.00161v1 Announce Type: new Abstract: Diffusion language models (DLMs) are an attractive alternative to autoregressive models because they promise sublinear-time, parallel generation, yet pr
Knowledge catalogue
arXiv:2605.00161v1 Announce Type: new Abstract: Diffusion language models (DLMs) are an attractive alternative to autoregressive models because they promise sublinear-time, parallel generation, yet pr
arXiv:2605.00421v1 Announce Type: new Abstract: Large language models (LLMs) show promise in radiology but their deployment is limited by computational requirements that preclude use in resource-const
arXiv:2603.17837v3 Announce Type: replace-cross Abstract: During conversational interactions, humans subconsciously engage in concurrent thinking while listening to a speaker. Although this internal c
arXiv:2604.27043v1 Announce Type: new Abstract: Today's AI assistants such as OpenClaw are designed to handle context effectively, making context learning an increasingly important capability for mode
arXiv:2512.14067v2 Announce Type: replace-cross Abstract: Diffusion language models (dLMs) have emerged as a promising paradigm that enables parallel, non-autoregressive generation, but their learning
arXiv:2604.28196v1 Announce Type: new Abstract: Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominant
arXiv:2604.27351v1 Announce Type: new Abstract: Agentic large language model systems have demonstrated strong capabilities. However, their reliance on language as the universal interface fundamentally
arXiv:2604.27093v1 Announce Type: cross Abstract: Current LLM safety alignment techniques improve model robustness against adversarial attacks, but overlook whether and how LLMs can recover helpfulnes
arXiv:2604.28185v1 Announce Type: new Abstract: Recent visual generation models have made major progress in photorealism, typography, instruction following, and interactive editing, yet they still str
arXiv:2604.27082v1 Announce Type: new Abstract: We present a framework for migrating production Large Language Model (LLM) based systems when the underlying model reaches end-of-life or requires repla
arXiv:2604.26478v1 Announce Type: new Abstract: Hyperspectral imaging (HSI) semantic segmentation typically relies on in-domain training, but limited data availability often restricts model performanc
Goblin and related magical mentions were overrewarded in training, and the behavior was reinforced over successive models. We removed the goblin-affine reward signal for future models, and filtered tr
Increasingly, I think, we will see a gap between what you can do with frontier model APIs & what you can do with the native apps from the frontier labs (Codex, Claude Code). Models developed and train
arXiv:2604.25934v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) as interactive agents has exposed a category of behavioral failure that prevailing terminology, princip
OpenAI: OpenAI says its models, starting with GPT-5.1, “increasingly mentioned goblins, gremlins, and other creatures”, leading to prompt instructions to mitigate it — Starting with GPT-5.1, our model
OpenAI is preparing to launch a new frontier cybersecurity model, GPT-5.5-Cyber. CEO Sam Altman said the model will not be available to the general public, but will be first rolled out to a select gro
arXiv:2505.24867v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have made impressive strides in understanding spatio-temporal relationships in videos. Howeve
$3/million output tokens. Qwen 3.5 Plus is basically a frontier model. Let that sink in. Introducing Qwen3.6-Plus from @Alibaba_Qwen, a 1M-context model built for real-world agents, agentic coding, an
arXiv:2604.25072v1 Announce Type: new Abstract: Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evalua
arXiv:2410.06723v2 Announce Type: replace-cross Abstract: Pathology foundation models (PFMs) have emerged as powerful pretrained encoders for computational pathology, but their robustness under clinic
arXiv:2604.25289v1 Announce Type: cross Abstract: Practically, training diffusion models typically requires explicit time conditioning to guide the network through the denoising sampling process. Espe
arXiv:2604.25200v1 Announce Type: cross Abstract: Public agencies are beginning to consider large language models (LLMs) as decision-support tools for grant evaluation. This creates a practical govern
arXiv:2508.18473v3 Announce Type: replace Abstract: While Large Language Models (LLMs) have emerged as powerful foundational models to solve a variety of tasks, they have also been shown to be prone t
arXiv:2604.25685v1 Announce Type: cross Abstract: Foundation segmentation models such as the Segment Anything Model (SAM) have demonstrated strong generalization across natural images; however, their
arXiv:2604.25781v1 Announce Type: new Abstract: Articulation modeling aims to infer movable parts and their motion parameters for a 3D object, enabling interactive animation, simulation, and shape edi
arXiv:2603.19124v2 Announce Type: replace Abstract: This paper presents the design, modeling, and fabrication of 3D-printed, tendon-actuated continuum robots featuring a flexible, tapered backbone con
arXiv:2604.23385v1 Announce Type: new Abstract: Transthoracic echocardiography is the reference standard for confirming structural heart disease (SHD), but first-line screening is limited by cost, wor
arXiv:2502.04424v4 Announce Type: replace-cross Abstract: With the integration of multimodal large language models (MLLMs) into robotic systems and AI applications, embedding emotional intelligence (E
feels like forever ago, but had to include openclaw in q1 trends coding models improved greatly in Q4 of 2025, early jan was ppl running claude codes in parallel, and clawdbot blew up late jan models
'If they ever tell my story let them say that I walked with giants'--Troy I am humbled&excited the model we released last week is trending #2 on @huggingface, between giant models such as DeepSeek,Qwe
arXiv:2604.23732v1 Announce Type: cross Abstract: Disease progression varies with age and is influenced by underlying genetic, biochemical, and hormonal etiologies, suggesting the need for tailored mo
arXiv:2603.01581v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models build a token-domain robot control paradigm, yet suffer from low speed. Speculative Decoding (SD) is an op
arXiv:2604.24542v1 Announce Type: cross Abstract: Large language models deployed at runtime can misbehave in ways that clean-data validation cannot anticipate: training-time backdoors lie dormant unti
arXiv:2604.24182v1 Announce Type: new Abstract: Current Vision-Language-Action (VLA) models predominantly rely on end-to-end fine-tuning. While effective, this paradigm compromises the inherent genera
OpenClaw 2026.4.26 🦞 🎙️ Google Live Talk 🦙 Better Ollama/local models 🧳 Bring over Claude + Hermes setups 🔐 One-command Matrix E2EE Big release. Local models eat well. https://github.com/openclaw/open
arXiv:2604.22823v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementar
Results on the remainder of our benchmarks are in for Kimi K2.6, the new #1 open-weight model on the Vals Index (#8 overall) Kimi K2.6 is competitive with many closed-weight models, at a fraction of t
arXiv:2511.07410v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) and Vision Language Models (VLMs) have become popular tools for embodied high-level planning. However, their depl
arXiv:2509.14837v2 Announce Type: replace Abstract: Recent advances in causal interpretability have extended from language models to vision-language models (VLMs), seeking to reveal their internal mec
arXiv:2510.08618v2 Announce Type: replace-cross Abstract: Omni-modal large language models (OLLMs) offer a promising end-to-end solution for slide-enhanced speech recognition due to their inherent mul
arXiv:2604.22748v1 Announce Type: new Abstract: As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a centra
arXiv:2602.01022v2 Announce Type: replace-cross Abstract: Behavioral parameters such as loss aversion, herding, and extrapolation are central to asset pricing models but remain difficult to measure re
Here's a uv one-liner that downloads and runs the MLX model against a local mp3 file uv run --with mlx-audio python -m mlx_audio.stt.generate --model mlx-community/VibeVoice-ASR-4bit --audio lenny.mp3
arXiv:2604.22266v1 Announce Type: new Abstract: Large Language Models often achieve strong performance by generating long intermediate chain-of-thought reasoning. However, it remains unclear when a mo
arXiv:2604.22633v1 Announce Type: cross Abstract: The Gaussian mixture model is widely used in unsupervised learning, owing to its simplicity and interpretability. However, a fundamental limitation of
arXiv:2604.22230v1 Announce Type: cross Abstract: Benchmark hacking refers to tuning a machine learning model to score highly on certain evaluation criteria without improving true generalization or fa
arXiv:2604.22464v1 Announce Type: new Abstract: Continual Model Merging (CMM) sequentially integrates task-specific models into a unified architecture without intensive retraining. However, existing C
No more updating needed for new model releases via Nous Portal and OpenRouter! Hermes will no longer have to be updated to receive model list curation updates for several providers, including Nous Por
arXiv:2507.15753v2 Announce Type: replace-cross Abstract: Generative machine learning models have revolutionized material discovery by capturing complex structure-property relationships, yet extending
API is Available Today! 🔹 Keep base_url, just update model to deepseek-v4-pro or deepseek-v4-flash. 🔹 Supports OpenAI ChatCompletions & Anthropic APIs. 🔹 Both models support 1M context & dual modes (T
arXiv:2604.21921v1 Announce Type: new Abstract: We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representati
deepseek-v4-flash is now available on Ollama's cloud! Hosted in the US. Try it with Claude Code: ollama launch claude --model deepseek-v4-flash:cloud Try it with OpenClaw: ollama launch openclaw --mod
arXiv:2604.20852v1 Announce Type: cross Abstract: Learning to rank (LTR) is one of the core tasks in Machine Learning. Traditional LTR models have made great progress, but nearly all of them are imple
arXiv:2508.19353v2 Announce Type: replace-cross Abstract: While transfer learning is an effective strategy, it often overlooks the opportunity to leverage knowledge from numerous available models onli
Simon Willison requested that Hugging Face add a repository size sorting option to help users find and filter models by storage requirements, specifically mentioning interest in locating DeepSeek quan
GPT-5.5 is a giant leap forward for handling ambiguity compared to previous GPT models. As Windsurf 2.0 focuses more on parallel agents, this model is key for long-horizon tasks — it excels at underst
arXiv:2604.21127v1 Announce Type: new Abstract: The NASA PACE mission provides unprecedented hyperspectral observations of ocean color, aerosols, and clouds, offering new insights into how these compo
arXiv:2604.21343v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) such as LLaVA are typically trained with an autoregressive language modeling objective, providing only indirect supervisi
arXiv:2507.01829v2 Announce Type: replace-cross Abstract: Multi-timescale sequence modeling relies on capturing both local fast dynamics and global slow context; yet, maintaining these capabilities un
arXiv:2604.21017v1 Announce Type: cross Abstract: Autonomous medical robots hold promise to improve patient outcomes, reduce provider workload, democratize access to care, and enable superhuman precis