TechniqueRLHF / Alignment8 recent entries11 Aug 2026LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning RoutingarXiv:2608.08889v1 Announce Type: new Abstract: Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large L→11 Aug 2026LIBAD: A Multimodal Anomaly Detection Benchmark for Li-Ion Battery Electrode ManufacturingarXiv:2608.07958v1 Announce Type: new Abstract: Multimodal industrial anomaly detection has largely focused on discrete products using strongly correlated RGB and 3D observations, leaving continuous p
TechniqueRAG8 recent entries28 Jul 2026Do Small Models Use the Law You Give Them? Context-Injected Fine-Tuning for Legal QA in BangladesharXiv:2607.23446v1 Announce Type: cross Abstract: A small language model can receive the governing statutory provision and still answer incorrectly. We test whether fine-tuning on examples containing →28 Jul 2026Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational ExcellencearXiv:2607.22948v1 Announce Type: cross Abstract: The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and →4 Aug 2026Seeing the Unseen: Towards Training-Free Inspection for Wind Turbine Blades Using Knowledge-Augmented Vision Language ModelsarXiv:2510.22868v2 Announce Type: replace Abstract: Wind turbine blades operate in harsh environments, making timely damage detection essential for preventing failures and optimizing maintenance. Dron→4 Aug 2026Douyin Multimodal Embedding Model Technical ReportarXiv:2608.02148v1 Announce Type: cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search→6 Aug 2026A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B TestingarXiv:2608.04625v1 Announce Type: new Abstract: Industrial recommendation strategy iteration heavily relies on large-scale A/B experimentation. Traditional tuning requires experts to repeatedly design→7 Aug 2026Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision SupportarXiv:2608.05151v1 Announce Type: cross Abstract: Wastewater operators need answers grounded in how their plant's variables interact and how fast effects propagate, not in generic pretraining text, wh→11 Aug 2026RAG-Based Auto-Configuration for Industrial Fieldbus DevicesarXiv:2608.08618v1 Announce Type: cross Abstract: Industrial device commissioning requires engineers to manually extract hundreds of protocol-specific parameters from heterogeneous PDF manuals and tra→12 Aug 2026Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy DatasetsarXiv:2608.10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing sin
TechniqueAgents8 recent entries11 Aug 2026A Communication-Efficient Digital Twin Framework for PSO-Based Swarm Navigation and Obstacle AvoidancearXiv:2406.19930v4 Announce Type: replace Abstract: Swarm-based target localization in industrial environments faces two major challenges: navigating obstacle-rich spaces and managing intensive commun→12 Aug 2026Twitch streamers can now opt out from training Amazon’s AITwitch users can now opt out of allowing their content to be used to train Amazon's generative AI models. Opting out means that 'your streams, VODs, clips, stream chats, and pictures and text on your →12 Aug 2026Recovering Wasted Compute in Autoresearch AgentsarXiv:2608.10424v1 Announce Type: new Abstract: A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such agents have in→12 Aug 2026Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT SystemsarXiv:2607.10309v2 Announce Type: replace Abstract: Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT→12 Aug 2026Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical CollaborationarXiv:2608.11195v1 Announce Type: new Abstract: AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case→12 Aug 2026LLM Agents Factory: Retrieval of Domain-Specific LLM AgentsarXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deploymen→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions→12 Aug 2026A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona ProblemarXiv:2608.10760v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has become the de-facto interface for connecting LLM agents to enterprise tools, and adoption has been explosive: wit
TechniqueFine-tuning8 recent entries11 Aug 2026SuperCoder: Assembly Program Superoptimization with Large Language ModelsarXiv:2505.11480v4 Announce Type: replace-cross Abstract: Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its inp→11 Aug 2026iLTM: Integrated Large Tabular ModelarXiv:2511.15941v2 Announce Type: replace-cross Abstract: Tabular data underpins decisions across science, industry, and public services. Despite rapid progress, advances in deep learning have not ful→12 Aug 2026VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video ForensicsarXiv:2608.11201v1 Announce Type: new Abstract: Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and auth→12 Aug 2026Stochastic Emulation of a Fully Coupled Preindustrial E3SMv3 SimulationarXiv:2608.10277v1 Announce Type: cross Abstract: We present a stochastic coupled emulator of E3SM version 3, built on the SamudrACE framework, which couples an atmosphere emulator (ACE2) with a full-→12 Aug 2026Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy DatasetsarXiv:2608.10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing sin→12 Aug 2026Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT SystemsarXiv:2607.10309v2 Announce Type: replace Abstract: Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT→12 Aug 2026LLM Agents Factory: Retrieval of Domain-Specific LLM AgentsarXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deploymen→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions
TechniqueMultimodal8 recent entries10 Aug 2026Shape Your Feed: An LLM-based Agentic System for Conversational RecommendationarXiv:2608.06632v1 Announce Type: new Abstract: Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., cl→10 Aug 2026How Malachyte solves retail’s cold-start problem with managed real-time AIWhat’s the best way to recommend products to little-known users? We’ve spent our careers trying to solve this problem for major companies like Spotify and Priceline, and it’s why Sidd founded Malachyt→10 Aug 2026FedVAR: Prototype-Aligned Federated Framework for Video Anomaly RecognitionarXiv:2608.06876v1 Announce Type: cross Abstract: In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intellig→11 Aug 2026SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident KnowledgearXiv:2608.09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance, →11 Aug 2026LIBAD: A Multimodal Anomaly Detection Benchmark for Li-Ion Battery Electrode ManufacturingarXiv:2608.07958v1 Announce Type: new Abstract: Multimodal industrial anomaly detection has largely focused on discrete products using strongly correlated RGB and 3D observations, leaving continuous p→11 Aug 2026ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly DetectionarXiv:2608.09789v1 Announce Type: new Abstract: Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) ca→12 Aug 2026VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video ForensicsarXiv:2608.11201v1 Announce Type: new Abstract: Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and auth→12 Aug 2026ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist DeferralarXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imagin
TechniqueSafety8 recent entries12 Aug 2026Rethinking Data Efficiency in Industrial Dense Prediction: Pretraining Coherence, Not Inductive Bias, Determines ViTs Low-Data AdvantagearXiv:2608.10590v1 Announce Type: new Abstract: Vision Transformers (ViTs) are widely believed to require more labeled data than CNNs for industrial dense prediction. Through controlled experiments on→12 Aug 2026Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature ReviewarXiv:2608.10047v1 Announce Type: cross Abstract: In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management (PHM). Machine Learning (ML) has→12 Aug 2026Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT SystemsarXiv:2607.10309v2 Announce Type: replace Abstract: Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT→12 Aug 2026MARCO: Click-Intent Decomposition for Calibrated Ads Conversion PredictionarXiv:2608.10562v1 Announce Type: new Abstract: Not all clicks are equal. Industrial ads ranking decouples conversion probability into click-through rate (CTR) and post-click conversion rate (CVR), ye→12 Aug 2026From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and RecommendationarXiv:2608.10182v1 Announce Type: cross Abstract: Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation. When the busines→12 Aug 2026FedCGR: Federated Cross-Domain Generative RecommendationarXiv:2608.10929v1 Announce Type: new Abstract: Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignment difficult→12 Aug 2026ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist DeferralarXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imagin→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions