TechniqueRLHF / Alignment8 recent entries11 Aug 2026Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation BudgetsarXiv:2608.08189v1 Announce Type: new Abstract: LLM-driven program discovery relies on rapid evaluator feedback, but many scientific and engineering tasks require high-fidelity simulations, hardware e→12 Aug 2026Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGIarXiv:2608.10730v1 Announce Type: cross Abstract: The pursuit of artificial general intelligence (AGI) rests on a seemingly self-evident premise: that general intelligence, the kind of flexible, domai
TechniqueRAG8 recent entries10 Aug 2026Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation ToolsarXiv:2608.07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI app→10 Aug 2026SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine OptimizationarXiv:2602.12187v2 Announce Type: replace-cross Abstract: Search-Augmented Generative Engines (SAGE) have emerged as a new paradigm for information access, bridging web-scale retrieval with generative→10 Aug 2026Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image GenerationarXiv:2608.06751v1 Announce Type: cross Abstract: Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical→11 Aug 2026RAG-Based Auto-Configuration for Industrial Fieldbus DevicesarXiv:2608.08618v1 Announce Type: cross Abstract: Industrial device commissioning requires engineers to manually extract hundreds of protocol-specific parameters from heterogeneous PDF manuals and tra→12 Aug 2026TRACE: Trustworthy Retrieval-Augmented Conversational EnginearXiv:2608.10176v1 Announce Type: new Abstract: Public service chatbots are expected to deliver recommendations from an underlying public service directory, while also making sure that the recommendat→12 Aug 2026Rethinking Text-Based Image Retrieval in Specific DomainarXiv:2608.10524v1 Announce Type: cross Abstract: Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, exis→12 Aug 2026RAG for regular users?One of the reasons I got into local LLMs was the possibility of getting answers using my own documents and books (a few hundreds) instead of having to search through them manually. However since I'm n→12 Aug 2026MIRA: Medical Image Reflection for Agentic DiagnosisarXiv:2608.10827v1 Announce Type: cross Abstract: Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading e
TechniqueAgents8 recent entries12 Aug 2026DuplexWorld: Can voice agents help you get through the day?arXiv:2608.10716v1 Announce Type: cross Abstract: Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing →12 Aug 2026Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction HistoriesarXiv:2608.10319v1 Announce Type: cross Abstract: Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering tasks. As devel→12 Aug 2026Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented ReasoningarXiv:2608.10438v1 Announce Type: new Abstract: Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. Fo→12 Aug 2026Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human DesignarXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, s→12 Aug 2026Blacksmith raises $45M to aid AI code validation as agentic development growsBlacksmith Software Inc. today announced it has raised 45 million in new funding for its continuous integration service, which combines code development with cloud-based testing instead of on the deve→12 Aug 2026Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn InteractionarXiv:2608.10239v1 Announce Type: new Abstract: Generative AI makes social-engineering attacks more fluent, adaptive, and scalable, increasing the need for LLM-based de- fenders that can protect users→12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent HarnessesarXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these →12 Aug 2026Apexon targets stalled AI pilots with three AgentRise additionsSanta Clara-based technology services firm Apexon Inc. today expanded AgentRise, its agentic artificial intelligence platform, with three new components. The additions are named AgentRise Polaris, Age
TechniqueFine-tuning8 recent entries11 Aug 2026From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource SettingsarXiv:2608.08896v1 Announce Type: new Abstract: Imaging device downtime is a major barrier to healthcare delivery in low- and middle-income countries (LMICs), often driven by limited access to special→11 Aug 2026FailForge: Distilling Procedural Competence from Persistent Failures into Code AgentsarXiv:2608.08570v1 Announce Type: new Abstract: Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering tasks, retaining →11 Aug 2026DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO…DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO, then deploy the fine-tuned model on Together AI for produc→11 Aug 2026Contamination Means Overestimation? A Fine-Grained Empirical Study in Code IntelligencearXiv:2506.02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread →12 Aug 2026Rethinking Text-Based Image Retrieval in Specific DomainarXiv:2608.10524v1 Announce Type: cross Abstract: Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, exis→12 Aug 2026RAG for regular users?One of the reasons I got into local LLMs was the possibility of getting answers using my own documents and books (a few hundreds) instead of having to search through them manually. However since I'm n→12 Aug 2026MIRA: Medical Image Reflection for Agentic DiagnosisarXiv:2608.10827v1 Announce Type: cross Abstract: Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading e→12 Aug 2026LLM Agents Factory: Retrieval of Domain-Specific LLM AgentsarXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deploymen
TechniqueMultimodal8 recent entries11 Aug 2026Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision-Language Models for Centimeter-Level Flood Depth EstimationarXiv:2608.07562v1 Announce Type: new Abstract: Urban flooding poses an escalating threat to transportation infrastructure, yet no operational system provides real-time, street-level flood-depth estim→11 Aug 2026Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended CitiesarXiv:2608.08045v1 Announce Type: new Abstract: Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Sim→11 Aug 2026I put Gemma 4 E4B and E2B into an e-reader so I can ask my weird questions and share my thoughts in private directly in app.Here's how it works in the app: Framework: Runs on LiteRT-LM (like Google's AI Edge). Models: Downloads either the E2B (~2.5 GB) or E4B (~3.6 GB) INT4 quantized models directly from ungated litert-com→11 Aug 2026I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examplesI wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon→11 Aug 2026CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics SimulationarXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, sup→12 Aug 2026Rethinking Text-Based Image Retrieval in Specific DomainarXiv:2608.10524v1 Announce Type: cross Abstract: Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, exis→12 Aug 2026Qwen3.8-2.4T-A95B is now live on Together AI. The Qwen Team’s latest flagship model is built for coding and long-horizon agent workflows, wi…The Qwen Team has released its flagship model, Qwen3.8‑2.4T‑A95B, on the Together AI platform (togethercompute) as of August 12 2026. This 2.4‑trillion‑parameter model is engineered for coding tasks a→12 Aug 2026DuplexWorld: Can voice agents help you get through the day?arXiv:2608.10716v1 Announce Type: cross Abstract: Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing
TechniqueSafety8 recent entries12 Aug 2026The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AIarXiv:2608.10153v1 Announce Type: new Abstract: Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSec→12 Aug 2026Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGIarXiv:2608.10730v1 Announce Type: cross Abstract: The pursuit of artificial general intelligence (AGI) rests on a seemingly self-evident premise: that general intelligence, the kind of flexible, domai→12 Aug 2026SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning AgentsarXiv:2608.10157v1 Announce Type: new Abstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over t→12 Aug 2026Partially Observable Learning for Multi-Platform Dispatch OptimizationarXiv:2608.10897v1 Announce Type: new Abstract: Instant delivery platforms have become a critical component of urban logistics, increasingly relying on crowdsourced couriers to fulfill highly dynamic →12 Aug 2026Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language ModelsarXiv:2608.10405v1 Announce Type: cross Abstract: Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in signi→12 Aug 2026Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) InteroperabilityarXiv:2608.10300v1 Announce Type: new Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human re→12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent HarnessesarXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these →12 Aug 2026APCReg: Anatomical-Prior-Guided Coarse-to-Fine CBCT--IOS Registration via Multi-View Projection and Reliability-Controlled Residual CorrectionarXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp