Loading page…
With the same RL setup and budget, Figure 3(a) shows downstream masking producing higher training reward and checkpoint accuracy than random masking, while upstream masking performs worst throughout training. · CiteArk