Kwai Summary Attention Technical Report
arXiv:2604.24432v1 Announce Type: cross Abstract: Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic unde
Knowledge catalogue
arXiv:2604.24432v1 Announce Type: cross Abstract: Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic unde
arXiv:2505.11772v3 Announce Type: replace Abstract: We introduce LAMP (Local Attribution Mapping Probe), a method that shines light onto a black-box language model's decision surface and studies how r
arXiv:2604.23314v1 Announce Type: new Abstract: Segmentation is central to clinical diagnosis and monitoring, yet the reliability of modern foundation models in medical imaging still depends on the av
arXiv:2604.22795v1 Announce Type: cross Abstract: This study presents a multi-agent reinforcement learning (MARL) framework for load-constrained wind farm flow control (WFFC). While wake steering can
Local model users get a lot of boring-good fixes: @ollama context handling, thinking controls, timeouts, local auth, discovery, and OpenAI-compatible proxy behavior. https://docs.openclaw.ai/providers
LTX Desktop 1.0.5 is an open-source desktop app for generating videos with LTX models on Windows/Linux NVIDIA GPUs locally or via API . The version 1.0.5 release includes bug fixes such as LTX API ins
A discussion about AI agents that can remember information from one interaction to the next, even across separate sessions , exploring how to implement persistent memory systems for personal AI assist
Tuna-2 is a unified multimodal model that performs visual understanding and generation directly based on pixel embeddings, employing simple patch embedding layers to encode visual input without a VAE
MiMo-V2.5-Pro is a fully open-sourced Mixture-of-Experts language model with 1.02T total parameters and 42B active parameters , available under the MIT License for commercial use, training, and fine-t
Laguna XS 2 is a compact language model available through Ollama's model library, designed for efficient inference on resource-constrained devices. The model represents an update or variant in the Lag
Nemotron-3 is an open-source language model available through Ollama's model library, likely representing NVIDIA's Nemotron-3 family of instruction-tuned models designed for various NLP tasks. The mod
arXiv:2604.22755v1 Announce Type: cross Abstract: Reliable decision support in nuclear engineering requires traceable, domain-grounded knowledge retrieval, yet safety and risk analysis workflows remai
ComfyUI has announced the arrival of ComfyStudio, a new development or feature release for their platform. This announcement was shared via a Substack blog post and promoted through their official X (
arXiv:2508.01635v2 Announce Type: replace-cross Abstract: Microservice architectures enable scalable cloud-native applications; however, the distributed nature of these systems complicates the mainten
arXiv:2604.23442v1 Announce Type: new Abstract: Weeds compete with crops for light, water, and nutrients, reducing yield and crop quality. Efficient weed detection is essential for site-specific weed
arXiv:2604.22967v1 Announce Type: cross Abstract: Trust Region Bayesian Optimization (TuRBO) is an effective strategy for alleviating the curse of dimensionality in high-dimensional black-box optimiza
arXiv:2604.23296v1 Announce Type: cross Abstract: Aspect Sentiment Quad Prediction (ASQP) has seen significant advancements, largely driven by the powerful semantic understanding and generative capabi
arXiv:2603.20738v2 Announce Type: replace Abstract: Cross-subject EEG-to-image retrieval for visual decoding is challenged by subject shift and hubness in the embedding space, which distort similarity
arXiv:2602.04120v2 Announce Type: replace-cross Abstract: Though Explainable AI (XAI) has made significant advancements, its inclusion in edge and IoT systems is typically ad-hoc and inefficient. Most
arXiv:2604.22939v1 Announce Type: cross Abstract: While the next-token prediction (NTP) paradigm enables large language models (LLMs) to express their intrinsic knowledge, its sequential nature constr
The ComfyStudio Artist Residency Program is an application-based initiative that provides opportunities for artists to participate in a structured residency. Interested applicants can submit their inf
Mark Gurman / Bloomberg: Sources: Apple plans an AI overhaul for photo editing in iOS 27, including using on-device AI models to extend, enhance, and reframe photos — Apple Inc. is planning a major ov
arXiv:2402.11789v5 Announce Type: replace-cross Abstract: Anomaly localization in images -- identifying regions that deviate from normal patterns -- is vital in applications such as medical diagnosis
arXiv:2601.21048v2 Announce Type: replace Abstract: Unsupervised neural combinatorial optimization (NCO) enables learning powerful solvers without access to ground-truth solutions. Existing approaches
arXiv:2507.04503v2 Announce Type: replace Abstract: Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and const
arXiv:2604.22846v1 Announce Type: new Abstract: The expanding ecosystem of pathology foundation models has produced powerful but fragmented tile-level representations, limiting their use in clinical t
arXiv:2502.03669v3 Announce Type: replace-cross Abstract: AI methods, such as generative models and reinforcement learning, have recently been applied to combinatorial optimization (CO) problems, espe
arXiv:2604.22872v1 Announce Type: new Abstract: Autonomous vehicles (AVs) rely on real-time perception systems to understand road environments and ensure safe navigation. However, implementing reliabl
Z-Image-Omni-Base is a versatile foundation model capable of both generation and editing tasks , designed for use within the ComfyUI framework. It is released to unlock the full potential for communit
arXiv:2511.13211v2 Announce Type: replace Abstract: Despite recent advancements in 3D-text cross-modal alignment, existing state-of-the-art methods still struggle to align fine-grained textual semanti
arXiv:2604.22748v1 Announce Type: new Abstract: As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a centra
arXiv:2604.22199v1 Announce Type: cross Abstract: Autonomous robots operating in open environments need the ability to continuously handle tasks that are not covered by predefined local methods. Howev
Release b8943 of llama.cpp fixes missing exports in llama-common and refactors the common/debug module, moving abort_on_nan from a template parameter to a member of base_callback_data. The release als
Release b8944 was published on April 27, 2026 , representing a recent build of the llama.cpp project. Llama.cpp provides LLM inference in C/C++ for running large language models efficiently on consume
Release b8946 includes a fix to remove duplicate wo_s scale after build_attn for Qwen3 and LLaMA models , along with other improvements to the llama.cpp library. This is an intermediate build release
The search results do not contain specific information about release b8948 itself. Based on the context from llama.cpp releases, b8948 is one of the continuously updated build releases from the llama.
The search results don't contain specific details about the b8950 release. Based on the naming convention and context, b8950 is an intermediate build release of llama.cpp, a C/C++ implementation of LL
The search results don't contain specific details about release b8952. Based on the information available, I can provide this knowledge base entry: b8952 is a release in the llama.cpp project, an open
ComfyUI now supports HappyHorse 1.0, Alibaba’s cinematic video generation model. Highlights: - Cinematic aesthetics, 1080p / 15s multi-shot consistency - Audio-visual sync: lip-sync dialogue, ambiance
Creative studios are being asked to do something fundamentally different now. Not just great work, but foundational systems for experience, brand, and content that their clients can own, extend, and b
arXiv:2604.22331v1 Announce Type: new Abstract: This study analyses simulated and real-world implementations of depth-aware rover navigation, highlighting the transition from stereo vision to monocula
arXiv:2511.10571v2 Announce Type: replace Abstract: Hidden Markov Models (HMMs) are fundamental for modeling sequential data, yet learning their parameters from observations remains challenging. Class
This appears to be the second part of a post series showcasing example outputs from ComfyUI, likely demonstrating generated images or workflow results on the ComfyUI Twitter account. Without access to
arXiv:2604.21950v1 Announce Type: cross Abstract: Small language models (1-3B) are practical to run locally, but individually limited on harder code generation tasks. We ask whether composing them int
arXiv:2604.22166v1 Announce Type: new Abstract: While language models demonstrate sophisticated syntactic capabilities, the extent to which their internal mechanisms align with cross-constructional pr
arXiv:2604.22339v1 Announce Type: new Abstract: Handling the dynamic environments is a significant research challenge in Visual Simultaneous Localization and Mapping (SLAM). Recent research combines 3
FLUX 2 Dev is an open-weight, 32-billion-parameter AI model developed by Black Forest Labs for text-to-image generation and advanced image editing . The model is available in ComfyUI and Diffusers fra
arXiv:2604.22416v1 Announce Type: cross Abstract: Latent variables pose a fundamental challenge to causal discovery and inference. Conventional local methods focus on direct neighbors but fail to prov
arXiv:2604.22160v1 Announce Type: cross Abstract: Human visual perception offers valuable insights for understanding computational principles of motion-based scene interpretation. Humans robustly dete
Get started: 1. Update ComfyUI (or try it on Comfy Cloud: https://links.comfy.org/3Qqbi3F 2. Open Template Library → search HappyHorse 3. Adjust prompt/input image → Run More details: https://links.co
Here's a uv one-liner that downloads and runs the MLX model against a local mp3 file uv run --with mlx-audio python -m mlx_audio.stt.generate --model mlx-community/VibeVoice-ASR-4bit --audio lenny.mp3
arXiv:2604.22679v1 Announce Type: cross Abstract: The increasing adoption of AI systems in hiring has raised concerns about algorithmic bias and accountability, prompting regulatory responses includin
arXiv:2604.22081v1 Announce Type: new Abstract: Most reinforcement-learning (RL) controllers used in continuous control are architecturally centralized: observations are compressed into a single laten
The NVIDIA Tesla K80 is a legacy GPU that can run modern LLMs through Ollama with community-maintained patches, since official Ollama dropped support for CUDA Compute Capability 3.7 hardware. Each K80
Llamacpp is a tool for running large language models locally on personal computers, offering free access without relying on cloud services. The technology emphasizes speed, security, and computational
This post from Ollama highlights the collaborative nature of open-source software development, suggesting that the value of OSS projects lies not just in the code itself but in the community and relat
arXiv:2604.21957v1 Announce Type: cross Abstract: Recent works have demonstrated that attention-based transformer and large language model (LLM) architectures can achieve strong channel state predicti
Microsoft's MIT licensed VibeVoice speech-to-text model (think Whisper with speaker diarization) is really good - my notes on running the 5.71GB 4bit MLX conversion on an M5 MacBook, using about 60GB
Xiaomi released and open-sourced MiMo-V2.5-Pro, delivering significant improvements over its predecessor in agentic capabilities, complex software engineering, and long-horizon tasks. The model is an
Ornstein-Hermes-27B-MLX live https://huggingface.co/GestaltLabs/Ornstein-Hermes-3.6-27b-MLX-16bit https://huggingface.co/GestaltLabs/Ornstein-Hermes-3.6-27b-MLX-8bit https://huggingface.co/GestaltLabs