It has been established that training a box-based detector network can enhance the localization performance of weakly supervised and unsupervised methods. Moreover, we extend this understanding by demonstrating that these detectors can be utilized to improve the original network, paving the way for further advancements. To accomplish this, we train the detectors on top of the network output instead of the image data and apply suitable loss backpropagation. Our findings reveal a significant improvement in phrase grounding for the ``what is where by looking'' task, as well as various methods of unsupervised object discovery. Our code is available at https://github.com/eyalgomel/box-based-refinement.
翻译:已有研究表明,训练基于框的检测器网络能够增强弱监督与无监督方法的定位性能。此外,我们通过证明此类检测器可用于改进原始网络,进一步拓展了这一认知,为后续发展铺平了道路。为实现这一目标,我们在网络输出之上而非图像数据上训练检测器,并应用适当的损失反向传播。实验结果显示,在“通过观察定位目标”任务中的短语定位以及多种无监督目标发现方法上均取得了显著提升。我们的代码开源地址为 https://github.com/eyalgomel/box-based-refinement。