nova-pro-v1 — LiveBench Scores
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
View sourceAmazon
Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December...
37.4
Quality Score
1260
Arena ELO
Undisclosed
Parameters
300K
Context
Sign in to join the discussion
0
Downloads
0
Likes
Dec 2024
Released
Benchmarks
20
Research
3
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
View sourcelanguage: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
View sourcelanguage: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
View sourcelanguage: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
View sourceSparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes over its full expert bank. We study a deployment problem: convert that checkpoint to a smaller standard MoE under an explicit expert budget, without adding a compression-specific online module. To address this, we introduce UniMoMo, a post-training compression framework formulated as a constrained graph coarsening problem. Rather than relying on parameter distance, UniMoMo groups experts based on their functional similarity, using an unlabeled calibration set to measure how similarly experts respond to shared recommendation states. To prevent performance degradation, we introduce a layer-adaptive protection mechanism that restricts the merging of high-traffic experts based on their routing exposure. Across Amazon Beauty, KuaiRec, and TenRec with 2, 4, and 6 MoE blocks, the final four-expert checkpoints obtain source-relative five-run mean NDCG@10 ratios of 99.92%--102.30% and measured A100 speedups of 1.28times--1.63times. An aggressive two-expert, top-1 operating point obtains ratios of 98.36%--104.24% and speedups of 1.47times--2.21times. These endpoint results evaluate the complete conversion-and-adaptation workflow and show that a trained recommendation MoE can be exported at multiple serving budgets.
Learning robust representations of authorial style is crucial for authorship attribution and AI-generated text detection. However, existing methods often struggle with content-style entanglement, where models learn spurious correlations between authors' writing styles and topics, leading to poor generalization across domains. To address this challenge, we propose Explainable Authorship Variational Autoencoder (EAVAE), a novel framework that explicitly disentangles style from content through architectural separation-by-design. EAVAE first pretrains style encoders using supervised contrastive learning on diverse authorship data, then finetunes with a Variational Autoencoder (VEA) architecture using separate encoders for style and content representations. Disentanglement is enforced through a novel discriminator that not only distinguishes whether pairs of style/content representations belong to the same or different authors/content sources, but also generates natural language explanation for their decision, simultaneously mitigating confounding information and enhancing interpretability. Extensive experiments demonstrate the effectiveness of EAVAE. On authorship attribution, we achieve state-of-the-art performance on various datasets, including Amazon Reviews, PAN21, and HRS. For AI-generated text detection, EAVAE excels in few-shot learning over the M4 dataset. Code and data repositories are available onlinehttps://github.com/hieum98/avae https://huggingface.co/collections/Hieuman/document-level-authorship-datasets.
RAG typically assumes centralized access to documents, which breaks down when knowledge is distributed across private data silos. We propose a secure Federated RAG system built using Flower that performs local silo retrieval, while server-side aggregation and text generation run inside an attested, confidential compute environment, enabling confidential remote LLM inference even in the presence of honest-but-curious or compromised servers. We also propose a cascading inference approach that incorporates a non-confidential third-party model (e.g., Amazon Nova) as auxiliary context without weakening confidentiality.
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7
language: 0.5 | coding: 0.5 | instruction_following: 1.0 | Overall: 0.7