LTX Color Shifting
The post discusses a color shifting issue observed when using LTX model in Stable Diffusion, where generated images exhibit unwanted color variations or drift during processing. Community members like
Knowledge catalogue
The post discusses a color shifting issue observed when using LTX model in Stable Diffusion, where generated images exhibit unwanted color variations or drift during processing. Community members like
This guide provides instructions for setting up local AI capabilities on a 2019 Mac Pro running Ubuntu 24.04, covering the installation and configuration of ROCm 7.2.3 GPU drivers, PyTorch 2.10 for ma
arXiv:2605.19792v1 Announce Type: new Abstract: Visually-grounded language models (VLMs) are highly effective in linking visual and textual information, yet they often struggle with basic classificati
arXiv:2605.18823v1 Announce Type: new Abstract: Digital twins (DTs) for urban transportation systems have gained increasing attention; however, their systematic evaluation in safety-critical scenarios
arXiv:2512.04452v2 Announce Type: replace-cross Abstract: NORi is a machine learning (ML) parameterization of ocean boundary layer turbulence that is physics-based and augmented with neural networks.
arXiv:2605.18930v1 Announce Type: cross Abstract: Memory-augmented large language model (LLM) agents use iterative reflection and self-evolution to solve complex tasks, but these mechanisms introduce
@ollama + @deepseek_ai v4 pro handled entire monthly dev reports on Eigent. github prs → word doc → slack message → sent to product-release channel. in just one prompt. fully local. the full walkthrou
arXiv:2605.19201v1 Announce Type: cross Abstract: Deep learning models detect pneumonia from chest X-rays with high accuracy, but the performance declines under domain shifts caused by differences in
arXiv:2604.01341v2 Announce Type: replace Abstract: Mathematical modeling of visual textures traces back to Julesz's intuition that texture perception in humans is based on local correlations between
arXiv:2605.19610v1 Announce Type: cross Abstract: We investigate the asymptotic properties of the Levy Adaptive B-spline (LABS) regression model, a Bayesian nonparametric method that incorporates B-sp
arXiv:2605.19257v1 Announce Type: new Abstract: Monocular SLAM historically suffers from scale ambiguity and tracking failure in dynamic environments. While recent vision foundation models (VFMs) prov
arXiv:2605.18763v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly applied to analyzing wearable sensing data, which are long-term, multimodal, and highly personalized. A
arXiv:2605.19928v1 Announce Type: cross Abstract: Counterfactual Regret Minimization (CFR) is the dominant algorithmic family for solving large imperfect-information games, underpinning breakthroughs
arXiv:2605.20158v1 Announce Type: cross Abstract: Large Vision Language Models (LVLMs) show promise in medical applications, but their inability to faithfully ground responses in visual evidence raise
arXiv:2505.17726v3 Announce Type: replace-cross Abstract: Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular,
arXiv:2605.19856v1 Announce Type: cross Abstract: Training very deep neural networks requires controlling the propagation of magnitudes across depth. Without such control, activations and gradients ma
arXiv:2605.18851v1 Announce Type: new Abstract: Recent advances in Reinforcement Learning (RL) have underscored its potential for incentivizing reasoning capabilities of Large Language Models (LLMs).
arXiv:2605.19727v1 Announce Type: new Abstract: Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by com
This article examines key emerging trends that are transforming AI infrastructure, likely covering topics such as distributed computing, edge AI deployment, GPU optimization, and evolving cloud infras
arXiv:2605.19491v1 Announce Type: new Abstract: Traditional whole slide image (WSI) analysis methods typically rely on the multiple instance learning (MIL) paradigm, which extracts patch-level feature
arXiv:2602.03839v2 Announce Type: replace Abstract: Bandwidth-constrained distributed reinforcement learning (RL) post-training of large language models is bottlenecked by two channels: weight synchro
arXiv:2605.20120v1 Announce Type: new Abstract: AI-assisted theorem proving can now generate substantial Lean developments for olympiad-level mathematics, but the evidential status of such development
Seedance 2.0 is a ComfyUI tool that transforms basic performance footage into cinematic-quality scenes through an automated workflow, regardless of the original shooting location. The technique appear
What hardware actually powers open-source AI? Not benchmarks. Not vendor marketing. Real-world community usage. We’re launching @huggingface Hardware: → trending GPUs & CPUs → VRAM distribution → infe
LetsEnhance with its Digital Art model is considered best for quality, delivering sharp lines and clearest eye detail with the highest output resolution ceiling. For a free desktop option, Upscayl off
arXiv:2605.19969v1 Announce Type: new Abstract: Decentralized learning (DL) is an emerging machine learning paradigm where nodes collaboratively train models without a central server. However, the col
arXiv:2605.17923v1 Announce Type: cross Abstract: In video generation models, particularly world models, training large-scale video diffusion Transformers (such as DiT and MMDiT) poses significant com
arXiv:2601.09722v2 Announce Type: replace-cross Abstract: In this work, we present an annotation framework that demonstrates how a multilingual LLM pretrained on a large corpus can be used as a teache
Agentmw is an open-source middleware framework designed to enhance AI agent reliability and efficiency across different models and frameworks. It addresses key operational challenges including mid-run
arXiv:2511.06316v3 Announce Type: replace Abstract: In low- and middle-income countries, public safety and urban planning initiatives frequently face a critical shortage of accurate, location-specific
arXiv:2605.16309v1 Announce Type: new Abstract: LLM-based agents can recover from individual execution errors, yet they repeatedly fail on the same fault when the underlying process knowledge--operato
arXiv:2605.16649v1 Announce Type: new Abstract: Recent diffusion-based video generators have achieved remarkable visual fidelity and prompt controllability, yet scaling them to ultra-high-resolution (
arXiv:2605.16579v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion is a powerful paradigm for streaming and interactive video generation. However, its reliance on softmax self-attenti
arXiv:2605.18748v1 Announce Type: new Abstract: Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and ref
arXiv:2605.16768v1 Announce Type: new Abstract: Semantic segmentation of multi-source remote sensing images is a fundamental task for Earth observation applications. Existing methods often struggle wi
B9222 is a llama.cpp release that adds support for the TRI (Triangle) operation in the Hexagon HTP backend with HVX kernel additions . The release includes optimizations for Hexagon hardware accelerat
arXiv:2605.18374v1 Announce Type: cross Abstract: Large language models (LLMs) typically approach combinatorial optimization as an inference-time procedure, solving each instance separately through sa
arXiv:2605.18194v1 Announce Type: new Abstract: While Multi-Modal Large Language Models (MLLMs) demonstrate impressive capabilities in general reasoning, their embodied spatial intelligence remains ha
arXiv:2605.17965v1 Announce Type: cross Abstract: Bug localization remains a key bottleneck in downstream software maintenance tasks, including root cause analysis, triage, and automated program repai
arXiv:2603.22570v2 Announce Type: replace Abstract: Active computer vision promises efficient, biologically plausible perception through sequential, localized glimpses, but lacks scalable general-purp
arXiv:2603.13652v2 Announce Type: replace Abstract: Attribution methods for Vision Transformers (ViTs) aim to identify image regions that influence model predictions, but producing faithful and well-l
arXiv:2605.16377v1 Announce Type: cross Abstract: Transparent and standardized reporting is essential for reproducible scientific research, yet adherence to reporting guidelines remains inconsistent b
arXiv:2605.18680v1 Announce Type: new Abstract: Metaverse platforms rely on creator-driven marketplaces where avatars are assembled from discrete, taxonomy-labeled 3D assets (e.g., tops, bottoms, shoe
arXiv:2605.16458v1 Announce Type: cross Abstract: Image restoration models are increasingly applied to degraded medical scans, but in safety-sensitive settings they must improve image quality without
arXiv:2605.16829v1 Announce Type: new Abstract: Discrete diffusion models are a powerful, emerging paradigm for code generation. They construct programs through iterative refinement of partially corru
arXiv:2605.17826v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) excel at multimodal reasoning, yet it remains unclear whether their answers are grounded in visual evidence or driven by
arXiv:2605.16373v1 Announce Type: cross Abstract: Early and accurate diagnosis and lesion localization of bone infections are crucial for clinical treatment. PET-CT integrates anatomical information f
arXiv:2605.18395v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit systematic political biases in voter simulations, but their underlying mechanisms and cross-lingual generalizatio
arXiv:2605.16742v1 Announce Type: new Abstract: Cortical surface registration is often driven by local geometric descriptors (e.g., sulcal depth and curvature). While this approach achieves geometric
arXiv:2605.16433v1 Announce Type: cross Abstract: Task allocation in smart manufacturing systems needs to operate under decentralized decision-making, dynamic workloads, and shared resource constraint
arXiv:2605.16797v1 Announce Type: new Abstract: Egocentric video is increasingly used as a data source for robot learning, activity understanding, and embodied AI research, but collecting it at scale
arXiv:2601.07122v2 Announce Type: replace-cross Abstract: While virtualization and resource pooling empower cloud networks with structural flexibility and elastic scalability, they inevitably expand t
arXiv:2605.17447v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have shown strong promise on Optical Character Recognition (OCR), yet the sheer number of visual tokens required to enco
arXiv:2605.18020v1 Announce Type: new Abstract: Federated Learning (FL) algorithms implicitly assume that clients passively comply with server-side orchestration by sharing local model updates upon se
arXiv:2605.18028v1 Announce Type: cross Abstract: Federated fine-tuning of Large Language Models faces severe statistical heterogeneity. However, existing model-level defenses often overlook the root
arXiv:2605.17231v1 Announce Type: cross Abstract: Activation steering methods modify intermediate representations of language models to control output behavior, but universally assume the activation s
arXiv:2605.17759v1 Announce Type: new Abstract: To circumvent the inherent fidelity bottlenecks and optimization misalignment of VAE-based latent diffusion, pixel-space diffusion models have emerged a
arXiv:2604.01404v2 Announce Type: replace-cross Abstract: How do language models retrieve entity-specific facts from their parameters? We investigate this question by searching for sparse, entity-sele
From months of work at ILM to a single compositor in ComfyUI. That's not a workflow change. That's a shift in who gets to do this work. The @ActionVFX team just launched Advanced AI Workflows for VFX
arXiv:2605.18271v1 Announce Type: cross Abstract: With the rapid emergence of personal AI agents based on Large Language Models (LLMs), implementing them on-device has become essential for privacy and