AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

Outsmarting the Chameleon: Counterfactual Decoupling for Tactical OOD Shifts in Live Streaming Risk Assessment

DGX agent

arXiv:2606.02946v1 Announce Type: new Abstract: Live streaming has emerged as a primary medium for social interaction and digital commerce, yet it is increasingly plagued by sophisticated risks. A fun

model-releasesarxiv-cs-lg
3 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

DGX agent

arXiv:2606.03890v1 Announce Type: new Abstract: Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

PatchScene: Patch-based Voxel Diffusion for Large-Scale Scene Completion

DGX agent

arXiv:2606.03915v1 Announce Type: new Abstract: We propose PatchScene, a novel diffusion-based framework for large-scale LiDAR scene completion. Unlike existing methods that rely on global latent repr

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents

DGX agent

arXiv:2606.03236v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have substantially advanced mobile agents, yet proactive mobile assistance remains challenging because agents m

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

PerchRL: Vision-Based Agile Perching on Inclined Platforms under Rapid and Irregular Motion

DGX agent

arXiv:2606.03441v1 Announce Type: cross Abstract: Autonomous vision-based perching of quadrotors on moving inclined platforms is critical for air-ground collaboration but remains challenging due to th

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios

DGX agent

arXiv:2602.05302v3 Announce Type: replace Abstract: We present an in-depth evaluation of LLMs' ability to negotiate, a central business task requiring strategic reasoning, theory of mind, and economic

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

PINNfluence: Interpreting PINNs through Influence Functions

DGX agent

arXiv:2409.08958v3 Announce Type: replace-cross Abstract: Physics-informed neural networks (PINNs) have emerged as a powerful deep learning approach for solving partial differential equations (PDEs) i

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Plan, Verify and Fill: A Structured Parallel Decoding Approach for Diffusion Language Models

DGX agent

arXiv:2601.12247v3 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) present a promising non-sequential paradigm for text generation, distinct from standard autoregressive (AR) a

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Plan2Map: A Multimodal Benchmark for Document-Grounded Geospatial Boundary Reconstruction from Planning Records

DGX agent

arXiv:2606.02747v1 Announce Type: cross Abstract: Planning records define restrictions over geographic areas, but their source documents often provide only indirect spatial evidence rather than machin

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Pretraining Language Models on Historical Text

DGX agent

arXiv:2606.02991v1 Announce Type: cross Abstract: We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires add

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Proof-Refactor: Refactoring Generated Formal Proofs into Modular Artifacts

DGX agent

arXiv:2606.03743v1 Announce Type: new Abstract: While Large Language Models (LLMs) have shown strong performance in generating formal proofs, their outputs often remain less readable, modular, maintai

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

ProtocolBench: Which LLM MultiAgent Protocol to Choose?

DGX agent

arXiv:2510.17149v3 Announce Type: replace Abstract: As large-scale multi-agent systems evolve, the communication protocol layer has become a critical yet under-evaluated factor shaping performance and

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Psi-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialogues

DGX agent

arXiv:2606.02754v1 Announce Type: new Abstract: Personalization is a crucial capability of modern language agents. However, current research primarily positions personalized agents as passive responde

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction

DGX agent

arXiv:2512.10888v3 Announce Type: replace Abstract: Table extraction (TE) is a key challenge in document understanding. Traditional approaches detect tables first, then recognize their structure. Rece

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models

DGX agent

arXiv:2606.03858v1 Announce Type: new Abstract: Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

q0: Primitives for Hyper-Epoch Pretraining

DGX agent

arXiv:2606.03938v1 Announce Type: cross Abstract: Multi-epoch training is becoming the standard now that compute is growing faster than the supply of high-quality text. But pretraining a single model

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Qift: Shift-Friendly No-Zero W2 Post-Training Quantization for Rotated W2A4/KV4 LLM Inference

DGX agent

arXiv:2606.02823v1 Announce Type: new Abstract: Two-bit weight quantization is attractive for memory-efficient LLM inference, but the standard W2 level set {-2,-1,0,+1} often collapses under aggressiv

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Quadratic integrate-and-fire neurons exhibit less fragmented loss landscapes and outperform leaky integrate-and-fire neurons in spike-based gradient descent

DGX agent

arXiv:2606.03935v1 Announce Type: cross Abstract: The ability to train spiking neural networks is essential for modeling biological neural networks as well as for neuromorphic computing. However, for

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

QUIVER: Quantum-Informed Views for Enhanced Representations in Large ML Models

DGX agent

arXiv:2606.02785v1 Announce Type: new Abstract: Large machine learning models benefit substantially from multimodal inputs that provide a complementary view of the same example. We introduce QUIVER (Q

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Qwen-Image-Flash: Beyond Objective Design

DGX agent

arXiv:2606.03746v1 Announce Type: cross Abstract: Few-step distillation has become an effective strategy for accelerating advanced visual generative models, yet prior work has largely focused on disti

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

RadarSFD: Single-Frame Diffusion with Pretrained Priors for Radar Point Clouds

DGX agent

arXiv:2509.18068v2 Announce Type: replace Abstract: Millimeter-wave radar provides robust perception in fog, smoke, dust, and low light, making it attractive for size-, weight-, and power-constrained

model-releasesarxiv-cs-ro
3 Jun 2026
Model Releases

Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA

DGX agent

arXiv:2606.03728v1 Announce Type: new Abstract: Retrieval-augmented generation systems for legal question answering typically retrieve passages based on semantic similarity and provide them to a langu

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions

DGX agent

arXiv:2606.03889v1 Announce Type: new Abstract: Agent benchmarks should reflect what users actually ask deployed agents to do, yet existing benchmarks often miss key realism properties of real develop

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

ReaLM: Residual Quantization Bridging Knowledge Graph Embeddings and Large Language Models

DGX agent

arXiv:2510.09711v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have recently emerged as a powerful paradigm for Knowledge Graph Completion (KGC), offering strong reasoning and

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Reasoning Structure of Large Language Models

DGX agent

arXiv:2606.03883v1 Announce Type: new Abstract: Large reasoning models (LRMs) are often evaluated using metrics such as final-answer accuracy or token count. However, identical scores on these metrics

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Reconstructing Objects along Hand Interaction Timelines in Egocentric Video

DGX agent

arXiv:2512.07394v2 Announce Type: replace Abstract: We introduce the task of Reconstructing Objects along Hand Interaction Timelines (ROHIT). We first define the Hand Interaction Timeline (HIT) from a

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Relational Linearity is a Predictor of Hallucinations

DGX agent

arXiv:2601.11429v2 Announce Type: replace-cross Abstract: Hallucination is a central failure mode of language models (LMs). We focus on hallucinations in response to questions like: 'Which instrument

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Reliability-Guided Depth Fusion for Glare-Resilient Navigation Costmaps

DGX agent

arXiv:2606.03421v1 Announce Type: new Abstract: Specular glare on reflective floors, glass boundaries, and glossy indoor surfaces frequently corrupts active-stereo RGB-D depth measurements, producing

model-releasesarxiv-cs-ro
3 Jun 2026
Model Releases

RESCAST-100K: A Comprehensive Dataset for Cross-Domain Residential Load and Indoor Temperature Forecasting

DGX agent

arXiv:2606.02852v1 Announce Type: new Abstract: Accurate short-term forecasting of residential energy load and indoor temperature is essential for home energy management systems, grid-level demand res

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Rethinking Molecular Text Representations for LLMs: An Empirical Study

DGX agent

arXiv:2606.03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use. We present a sys

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation

DGX agent

arXiv:2606.03784v1 Announce Type: new Abstract: Embodied chain-of-thought (CoT) aims to bridge linguistic reasoning and robotic control, but its effective form and integration strategy remain underexp

model-releasesarxiv-cs-ro
3 Jun 2026
Model Releases

RobotValues: Evaluating Household Robots When Human Values Conflict

DGX agent

arXiv:2606.03312v1 Announce Type: cross Abstract: While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations in which robo

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

ROBUST-WT: Robust Uncertainty-aware Segmentation Transform via Whitening and Training Enhancements

DGX agent

arXiv:2606.03069v1 Announce Type: cross Abstract: Generalized segmentation of medical images prevents performance degradation when different imaging devices and clinical protocols are used across mult

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

RogueMerge: Robust and Unified Attacks against LLM Model Merging

DGX agent

arXiv:2606.03344v1 Announce Type: cross Abstract: Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a cri

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

DGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series

DGX agent

arXiv:2606.03301v1 Announce Type: new Abstract: We introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Sample-Size Scaling of the African Languages NLI Evaluation

DGX agent

arXiv:2606.03219v1 Announce Type: new Abstract: African languages have very little labelled data, and it is unclear if augmenting the quantity of annotation data reliably enhances downstream performan

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Samudra 2: Scaling Ocean Emulators across Resolutions

DGX agent

arXiv:2606.02610v1 Announce Type: cross Abstract: Ocean general circulation models (OGCMs) are essential to climate science but computationally expensive, limiting ensemble size and forcing scenarios.

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Scalable On-Hardware Training of Quantum Neural Networks and Application to Clinical Data Imputation

DGX agent

arXiv:2606.03517v1 Announce Type: cross Abstract: Training quantum neural networks (QNNs) on quantum hardware is currently bottlenecked by the cost of gradient estimation: standard parameter-shift met

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

SCOPE: Real-Time Natural Language Camera Agent at the Edge

DGX agent

arXiv:2606.02951v1 Announce Type: cross Abstract: Deploying language-driven agents in robotics requires evaluations that reflect real-world task demands: natural-language instructions with reproducibl

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

ScoreStop: Gradient-based early stopping using functional score tests

DGX agent

arXiv:2606.02740v1 Announce Type: cross Abstract: Gradient boosted decision trees require a stopping rule to avoid overfitting. The standard rule monitors a validation loss and stops if the loss fails

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

scTranslation: A Comprehensive Benchmark for Single-Cell Multi-Omics Modality Translation

DGX agent

arXiv:2606.03906v1 Announce Type: new Abstract: Simultaneous measurement of multiple omics modalities in single cells enables researchers to gain a more comprehensive understanding of cellular states

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding

DGX agent

arXiv:2606.03284v1 Announce Type: new Abstract: Frontier LLMs perform well in Western contexts, but remain poorly tested on underrepresented cultures such as those in Southeast Asia (SEA). Existing NL

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

See, Infer, Intervene: Proactive World Modeling for Goal-Oriented Social Intelligence

DGX agent

arXiv:2606.03371v1 Announce Type: new Abstract: Multimodal retail agents should not only recognize what a customer is doing, but also decide whether and how to assist before an explicit request is mad

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

SeeTraceAct: Visibility-Aware Latent Planning from Cross-Embodiment Demonstration Videos

DGX agent

arXiv:2606.02745v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) are promising general-purpose robot policies, but adapting them to new tasks typically requires costly task-speci

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

SenseJudge: Human-Centric Preference-Driven Judgment Framework

DGX agent

arXiv:2606.03189v1 Announce Type: new Abstract: Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is becoming an increasingly accepted paradigm. However

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill

DGX agent

arXiv:2606.03980v1 Announce Type: cross Abstract: Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) p

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale

DGX agent

arXiv:2606.03056v1 Announce Type: new Abstract: As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: skills depend

model-releasesarxiv-cs-ai
3 Jun 2026
← Previous
1…161162163164165…361
Next →