With the advancement of machine translation, post-editing (PE) has become a dominant workflow, making the accurate assessment of post-machine translation editing competence (PMTE) critical. However, performance-based PMTE assessments are vulnerable to subjective rater effects, compromising their validity. This study employs the Many-Facet Rasch model (MFRM) to conduct an in-depth analysis of rater-criterion interaction, a complex bias in PMTE evaluation. The research involved 144 examinees translating a scientific text and four expert raters assessing the outputs using four criteria: Logical relations, completeness, terminology, and fluency. The MFRM analysis successfully calibrated examinee ability, demonstrating high reliability (.86), and revealed significant variations in rater severity and criterion difficulty. Critically, the analysis identified specific quality control issues, including inconsistent scoring by one rater and ambiguity in the “Terminology” criterion. The bias analysis uncovered significant rater-criterion interactions. These findings demonstrate that MFRM is a powerful diagnostic tool that transforms assessment from a purely evaluative act into a mechanism for data-driven improvement. It provides objective, actionable evidence for refining scoring rubrics and conducting targeted rater training, thereby enhancing the fairness and validity of PMTE assessment.