With the same RL setup and budget, Figure 3(a) shows downstream masking producing higher training reward and checkpoint accuracy than random masking, while upstream masking performs worst throughout training. · CiteArk
With the same RL setup and budget, Figure 3(a) shows downstream masking producing higher training reward and checkpoint accuracy than random masking, while upstream masking performs worst throughout training.
来源:paper-fixed:PDF pp. 7 and 9, Section 4.2 Q4 and Figure 3(a)