[2103.11886] DeepViT: Towards Deeper Vision Transformer
cost. The pro-posed method makes it feasible to train deeper ViT models with consistent performance improvements via minor modification to existing ViT models. Notably, when training a deep ViT model with 32 transformer blocks, the Top-1 classification accuracy can be improved by 1.6% on ImageNet. Code is publicly available at this https UR...
arxiv.org
没有更多结果了~
- 意见反馈
- 页面反馈