TechniqueRLHF / Alignment8 recent entries11 Aug 2026Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human ResourcesarXiv:2608.07477v1 Announce Type: cross Abstract: This thesis examines the fairness of Automated Machine Learning (AutoML) tools in human resource hiring systems through the combined lenses of regulat→11 Aug 2026Beyond cognacyarXiv:2507.03005v3 Announce Type: replace Abstract: Computational phylogenetics has become an established tool in historical linguistics, with many language families now analyzed using likelihood-base→
TechniqueRAG8 recent entries7 Aug 2026CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor SearcharXiv:2508.02091v4 Announce Type: replace-cross Abstract: Approximate nearest-neighbor search (ANNS) algorithms have become increasingly critical for recent AI applications, particularly in retrieval-→10 Aug 2026Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation ToolsarXiv:2608.07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI app→10 Aug 2026Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat IntelligencearXiv:2608.06778v1 Announce Type: cross Abstract: Mapping cyber threat intelligence (CTI) text to MITRE ATT&CK techniques is essential for structured threat analysis, yet manual annotation is costly a→10 Aug 2026Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report GenerationarXiv:2608.07117v1 Announce Type: new Abstract: Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for →11 Aug 2026RAG-Based Auto-Configuration for Industrial Fieldbus DevicesarXiv:2608.08618v1 Announce Type: cross Abstract: Industrial device commissioning requires engineers to manually extract hundreds of protocol-specific parameters from heterogeneous PDF manuals and tra→11 Aug 2026Guardian Crawler: Retrieval-First Knowledge Discovery with Bounded LLM Augmentation for Noisy Web IntelligencearXiv:2608.08994v1 Announce Type: cross Abstract: Retrieving relevant evidence from noisy web data is challenging, particularly in sensitive domains containing incomplete reports, heterogeneous langua→12 Aug 2026Order Matters: LVLMs as Judges for Temporal Reasoning in Image SequencesarXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud→12 Aug 2026EvoMem: Memory-Augmented Evolution for Code OptimizationarXiv:2608.10795v1 Announce Type: new Abstract: Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may tran
TechniqueAgents8 recent entries12 Aug 2026Blacksmith raises $45M to aid AI code validation as agentic development growsBlacksmith Software Inc. today announced it has raised 45 million in new funding for its continuous integration service, which combines code development with cloud-based testing instead of on the deve→12 Aug 2026Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn InteractionarXiv:2608.10239v1 Announce Type: new Abstract: Generative AI makes social-engineering attacks more fluent, adaptive, and scalable, increasing the need for LLM-based de- fenders that can protect users→12 Aug 2026Automating and Scaling Behavioral Scientific Research on AI AgentsarXiv:2608.10030v1 Announce Type: new Abstract: As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI→12 Aug 2026Apexon targets stalled AI pilots with three AgentRise additionsSanta Clara-based technology services firm Apexon Inc. today expanded AgentRise, its agentic artificial intelligence platform, with three new components. The additions are named AgentRise Polaris, Age→12 Aug 2026Agentic Instruction Data Selection: Let DataMaster Interpret Your IntentarXiv:2608.10579v1 Announce Type: new Abstract: Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes it impractica→12 Aug 2026Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using AgentsarXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa→12 Aug 2026A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona ProblemarXiv:2608.10760v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has become the de-facto interface for connecting LLM agents to enterprise tools, and adoption has been explosive: wit→12 Aug 202618 two-word AI prompts I'm kind of obsessed with: 1) now what - great for when you've wrapped up a project or big push and you still have en…18 two-word AI prompts I'm kind of obsessed with: 1) now what - great for when you've wrapped up a project or big push and you still have energy and want AI to give you more 2) plz fix - usually accom
TechniqueFine-tuning8 recent entries11 Aug 2026Contamination Means Overestimation? A Fine-Grained Empirical Study in Code IntelligencearXiv:2506.02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread →12 Aug 2026VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video ForensicsarXiv:2608.11201v1 Announce Type: new Abstract: Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and auth→12 Aug 2026TACTICL: Task-Aware Compression of Tabular ICL ModelsarXiv:2608.10837v1 Announce Type: cross Abstract: The strong performance of foundation models for tabular tasks comes at substantial inference costs. Distilling models into task-specific architectures→12 Aug 2026Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASRarXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-reso→12 Aug 2026***Lights, inference, action…*** I’m so happy to share that @sequoia has led the Seed in @previewio. Developers are flying in magical AI-nat…***Lights, inference, action…*** I’m so happy to share that @sequoia has led the Seed in @previewio. Developers are flying in magical AI-native editors. But creative tooling is still stuck in the pre-→12 Aug 2026Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-TrainingarXiv:2602.01747v2 Announce Type: replace Abstract: Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world setting→12 Aug 2026CHORUS: Complementary Experts for High-Coverage Testbench Stimulus GenerationarXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation al→12 Aug 2026Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate ShiftarXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a eta-fraction of shifted tar
TechniqueMultimodal8 recent entries11 Aug 2026From Visual to Multimodal: Systematic Ablation of Encoders and Fusion Strategies in Animal IdentificationarXiv:2603.02270v2 Announce Type: replace Abstract: Automated animal identification is a practical task for reuniting lost pets with their owners, yet current systems often struggle due to limited dat→11 Aug 2026Curriculum Generation under Structured Parametric Environments for Robust Navigation PoliciesarXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f→11 Aug 2026AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference OptimizationarXiv:2608.07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist en→12 Aug 2026VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video ForensicsarXiv:2608.11201v1 Announce Type: new Abstract: Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and auth→12 Aug 2026Order Matters: LVLMs as Judges for Temporal Reasoning in Image SequencesarXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud→12 Aug 2026HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language ModelsarXiv:2506.03922v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchma→12 Aug 2026Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasetsarXiv:2608.11076v1 Announce Type: new Abstract: Automated lesion segmentation in whole-body PET/CT imaging can assist clinicians with cancer detection, staging, and treatment planning across radiotrac→12 Aug 2026Evaluating Semantic and Spatial Guidance for Foundation Model Segmentation of Small-Scale PV in Remote Sensing ImageryarXiv:2608.10801v1 Announce Type: new Abstract: Spatio-temporal PV data are essential for understanding adoption processes in off-grid regions, yet such data remain largely unavailable. Automated segm
TechniqueSafety8 recent entries12 Aug 2026Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven ApproacharXiv:2602.16481v2 Announce Type: replace Abstract: Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the effects of→12 Aug 2026Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight RegimesarXiv:2608.10209v1 Announce Type: new Abstract: Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with hu→12 Aug 2026Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-TrainingarXiv:2602.01747v2 Announce Type: replace Abstract: Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world setting→12 Aug 2026Easy3D-Labels: Supervising Semantic Occupancy Estimation with 3D Pseudo-Labels for Automotive PerceptionarXiv:2509.26087v5 Announce Type: replace Abstract: In perception for automated vehicles, safety is critical not only for the driver but also for other agents in the scene, particularly vulnerable roa→12 Aug 2026Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate ShiftarXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a eta-fraction of shifted tar→12 Aug 2026Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online SafetyarXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis →12 Aug 2026APCReg: Anatomical-Prior-Guided Coarse-to-Fine CBCT--IOS Registration via Multi-View Projection and Reliability-Controlled Residual CorrectionarXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp→12 Aug 2026Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using AgentsarXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa