AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

Agent-Orchestrated Adaptive RAG: A Comparative Study on Structured and Multi-Hop Retrieval

DGX agent

arXiv:2606.05658v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding their responses in external knowledge, but conventional pipeli

model-releasesarxiv-cs-ai
6 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Agentic Monte Carlo: Simulating Reinforcement Learning for Black-Box Agents

DGX agent

arXiv:2606.05296v1 Announce Type: cross Abstract: LLM agents operate in two distinct regimes: open-weight agents amenable to reinforcement learning (RL) and black-box agents whose behaviour must be co

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

AttackPathGNN: Cross-function vulnerability detection in smart contracts using state interference graphs and conjunction pooling

DGX agent

arXiv:2606.05986v1 Announce Type: cross Abstract: Existing learning-based detectors for Solidity smart-contracts reduce vulnerability detection to syntactic pattern matching within single functions, y

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Benchmark Everything Everywhere All at Once

DGX agent

arXiv:2606.06462v1 Announce Type: new Abstract: Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their co

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Benchmarking Counterfactual Prediction in Epidemic Time Series with Time-Varying Interventions

DGX agent

arXiv:2606.05692v1 Announce Type: cross Abstract: Deep learning has enabled significant advances in time-series causal inference, yet progress remains constrained by the lack of realistic benchmarks w

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillatio

DGX agent

arXiv:2606.05682v1 Announce Type: new Abstract: Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost c

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Brick-Composer: Using MLLMs for Assembly with Diverse Bricks

DGX agent

arXiv:2606.05445v1 Announce Type: new Abstract: We dream of AI agents that can read arbitrary designs and construct real-world objects from reusable building blocks. As a first step toward this vision

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Can AI Refute Economic Theory? Evidence from Beyond the Knowledge Cutoff

DGX agent

arXiv:2606.05383v1 Announce Type: cross Abstract: Can artificial intelligence (AI) refute economic theory? I document experiments in which I asked several AI models (Gemini, Refine, Claude, and ChatGP

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Can LLMs Write Correct TLA+ Specifications? Evaluating Natural-Language-to-TLA+ Generation

DGX agent

arXiv:2606.05792v1 Announce Type: new Abstract: TLA+ has supported industrial verification at companies such as Amazon and Microsoft, yet writing correct TLA+ specifications from natural language stil

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications

DGX agent

arXiv:2512.15231v3 Announce Type: replace Abstract: The automated and intelligent processing of massive remote sensing (RS) datasets is critical in Earth observation (EO). Existing automated systems a

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs

DGX agent

arXiv:2606.05966v1 Announce Type: cross Abstract: Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) sti

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Closing the Loop on Latent Reasoning via Test-Time Reconstruction

DGX agent

arXiv:2606.06252v1 Announce Type: new Abstract: Recent work moves intermediate reasoning from natural-language traces into latent or cache-level representations to reduce token overhead and avoid a di

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model

DGX agent

arXiv:2606.06099v1 Announce Type: new Abstract: Whether Large Language Models (LLMs) exhibit covert psychological manipulation in complex human-AI interactions has garnered increasing safety concerns.

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Critic-Guided Heterogeneous Multi-Agent Reasoning for Reliable Mathematical Problem Solving

DGX agent

arXiv:2606.05704v1 Announce Type: new Abstract: Recent Large Language Models (LLMs) have shown impressive reasoning abilities; but they are still susceptible to hallucinations, intermediate reasoning

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Cross-Epoch Adaptive Rollout Optimization for RL Post-Training

DGX agent

arXiv:2606.05606v1 Announce Type: cross Abstract: LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed ro

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

CTIConnect: A Benchmark for Retrieval-Augmented LLMs over Heterogeneous Cyber Threat Intelligence

DGX agent

arXiv:2510.11974v2 Announce Type: replace-cross Abstract: Cyber Threat Intelligence (CTI) is foundational to modern cybersecurity, enabling organizations to proactively defend against evolving threats

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Data Flow Control: Data Safety Policies for AI Agents

DGX agent

arXiv:2606.05679v1 Announce Type: cross Abstract: Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users. While recent work improves query correctness,

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows

DGX agent

arXiv:2606.05670v1 Announce Type: new Abstract: Does adding more agents help an LLM workflow once compared systems share the same benchmark loader, tool access, answer contract, usage accounting, and

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

DPBench: Structural Determinants of Multi-Agent LLM Coordination Under Simultaneous Resource Contention

DGX agent

arXiv:2602.13255v2 Announce Type: replace Abstract: We present DPBench, a benchmark for evaluating coordination in multi-agent systems built from large language models. Existing benchmarks measure tas

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions

DGX agent

arXiv:2606.06322v1 Announce Type: new Abstract: GUI agents - vision-based models that control desktops, web browsers, and mobile devices through graphical user interfaces - promise to automate a wide

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing

DGX agent

arXiv:2606.05950v1 Announce Type: new Abstract: Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing methods remain con

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Enhancing Software Engineering Through Closed-Loop Memory Optimization

DGX agent

arXiv:2606.05646v1 Announce Type: cross Abstract: Large language models (LLMs) have enabled powerful software engineering (SE) agents capable of navigating complex codebases and resolving real-world i

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Evaluating Agentic Configuration Repair for Computer Networks

DGX agent

arXiv:2606.06212v1 Announce Type: new Abstract: Misconfigurations in computer networks remain a major source of critical Internet outages. Research is turning to Large Language Models (LLMs) to automa

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Evaluation of LLMs for Mathematical Formalization in Lean

DGX agent

arXiv:2606.05632v1 Announce Type: new Abstract: Within the past few years, the ability of Large Language Models (LLMs) to generate formal mathematical proofs has improved drastically. We provide a com

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Exploring LLMs for South Asian Music Understanding and Generation

DGX agent

arXiv:2606.05522v1 Announce Type: cross Abstract: Recent advancements in Large Language Models (LLMs) have shown promising results in music understanding and generation tasks. However, existing works

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

GenTI: Benchmarking LLMs for Autonomous IDPS Rule Generation for Unseen Attacks

DGX agent

arXiv:2606.05844v1 Announce Type: cross Abstract: Rule-based Intrusion Detection and Prevention Systems (IDPS) offer precise attack detection as well as mitigation, however their manually crafted, sig

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Geographic Bias and Diversity in AI Evaluation

DGX agent

arXiv:2606.05187v1 Announce Type: cross Abstract: Among the many challenges hindering the responsible development and deployment of AI, arguably none has faced more intense scrutiny than bias in its v

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

GITCO: Gated Inference-Time Context Optimization in TSFMs

DGX agent

arXiv:2606.05332v1 Announce Type: new Abstract: Patch-based Time Series Foundation Models (TSFMs) suffer from context poisoning: structurally anomalous patches capture disproportionate attention and s

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement

DGX agent

arXiv:2606.06468v1 Announce Type: new Abstract: We introduce Goedel-Architect, an agentic framework for formal theorem proving in Lean 4 centered on blueprint generation and refinement. A blueprint is

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection

DGX agent

arXiv:2606.05566v1 Announce Type: new Abstract: Large Language Models (LLMs) have transformed natural language processing, but they remain vulnerable to Prompt Injection (PI) and Jailbreak (JB) attack

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment

DGX agent

arXiv:2606.05256v1 Announce Type: new Abstract: This study analyzes a publicly released dataset from a discontinued field experiment on Reddit's r/ChangeMyView. The intervention, conducted by unknown,

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

I Know What You Meme, Even If it Emerged Today: Understanding Evolving Memes through Open-World Knowledge Acquisition

DGX agent

arXiv:2606.05316v1 Announce Type: new Abstract: Multimodal memes are dynamic and often require up to date background knowledge for interpretation. Existing methods often overlook such knowledge or rel

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Learning What Matters: Probabilistic Task Selection via Mutual Information for Model Finetuning

DGX agent

arXiv:2507.12612v3 Announce Type: replace-cross Abstract: Supervised fine-tuning performance for large language models depends strongly on how training budget is distributed across a heterogeneous set

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

LLMCodec: Adapting Video Codecs for Efficient Weight Compression of Large Language Models

DGX agent

arXiv:2606.05861v1 Announce Type: cross Abstract: The rapid development of large language models(LLMs) has led to remarkable advances in natural language processing. However, the increasing scale of t

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents

DGX agent

arXiv:2606.06036v1 Announce Type: new Abstract: Despite recent progress, LLM agents still struggle with reasoning over long interaction histories. While current memory-augmented agents rely on a stati

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models

DGX agent

arXiv:2606.05429v1 Announce Type: new Abstract: Post-training quantization (PTQ) is critical for the efficient deployment of large language models (LLMs). Recent ultra-low-bit PTQ methods rely on rigi

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Multilingual Fine-Tuning via Localized Gradient Conflict Resolution

DGX agent

arXiv:2606.05613v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) has established cross-lingual versatility as a defining feature of modern systems. However, fine-tun

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

No Need to Train Your RDB Foundation Model

DGX agent

arXiv:2602.13697v2 Announce Type: replace Abstract: Relational databases (RDBs) contain vast amounts of heterogeneous tabular information that can be exploited for predictive modeling purposes. But si

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

OPRD: On-Policy Representation Distillation

DGX agent

arXiv:2606.06021v1 Announce Type: cross Abstract: On-policy distillation (OPD) supervises the student only in output space by matching next-token probabilities. This output-only paradigm has two limit

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

PC Layer: Polynomial Weight Preconditioning for Improving LLM Pre-Training

DGX agent

arXiv:2606.06470v1 Announce Type: cross Abstract: We propose a preconditioning (PC) layer, a weight parameterization via polynomial preconditioner that ensures stable weight conditioning throughout LL

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows

DGX agent

arXiv:2605.12376v2 Announce Type: replace Abstract: Table processing-including cleaning, transformation, augmentation, and matching-is a foundational yet error-prone stage in real-world data pipelines

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

DGX agent

arXiv:2606.05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically perf

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Retry Policy Gradients in Continuous Action Spaces

DGX agent

arXiv:2606.05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that the

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing

DGX agent

arXiv:2605.04733v2 Announce Type: replace Abstract: Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Reward Learning through Ranking Mean Squared Error

DGX agent

arXiv:2601.09236v3 Announce Type: replace-cross Abstract: Reward design remains a significant bottleneck in applying reinforcement learning (RL) to real-world problems. A popular alternative is reward

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

DGX agent

arXiv:2606.05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and r

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

SagnacAssisted Enhanced OTDR for Distributed Acoustic Sensing: A Standardized Benchmark and Engineering Evaluation Framework

DGX agent

arXiv:2606.05754v1 Announce Type: cross Abstract: Phase-sensitive optical time-domain reflectometry (phi-OTDR) is widely used in large-scale distributed acoustic sensing (DAS) because it provides dist

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

SciVisAgentSkills: Design and Evaluation of Agent Skills for Scientific Data Analysis and Visualization

DGX agent

arXiv:2606.05525v1 Announce Type: new Abstract: Recent advances in agentic visualization have enabled the translation of natural language into executable scientific visualization (SciVis) workflows. W

model-releasesarxiv-cs-ai
6 Jun 2026
← Previous
1…150151152153154…361
Next →