TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
DGX agentarXiv:2603.22867v1 Announce Type: cross Abstract: Multimodal stacks that mix ViTs, CNNs, GNNs, and transformer NLP strain embedded platforms because their compute/memory patterns diverge and hard real