pytorch

mirror of https://github.com/zebrajr/pytorch.git synced 2025-12-07 12:21:27 +01:00

Author	SHA1	Message	Date
Aapo Kyrola	02f0c1c9d7	make memonger work with RecurrentNetwork(Gradient) Summary: This diff enables support of recurrent networks for memonger: 1. Memonger descends into the step-nets and renames the blobs accordingly 2. Memonger tells the gradient op about the renamed blobs by adding a parameter "paramname.renamed=<new name>" 3. RecurrentNetworkGradientOp applies remapping to links and gradient blobs. I first thought of refactoring the whole gradient blob management of the recurrent network, but that looks to be very hard without a major revise of the code. Note, I did not enable memonger for neural_mt, since I think the team should do more testing before enabling this. Reviewed By: salexspb Differential Revision: D4812823 fbshipit-source-id: 1ffdf3cfb4fcd00eec5bb0ece3bf416aa6d3e26b	2017-04-05 09:48:25 -07:00
Aaron Markham	58f7f2b441	doxygen python block added Summary: Closes https://github.com/caffe2/caffe2/pull/226 Differential Revision: D4793550 Pulled By: JoelMarcey fbshipit-source-id: cc33e58186304fa8dcac2ee9115dcc271d785b1e	2017-03-29 06:46:16 -07:00
Viswanath Sivakumar	9775ffc6ae	Fixes to topological sort, canonical blob naming, sharing final blob Summary: Three small changes: Reviewed By: ajtulloch Differential Revision: D4437131 fbshipit-source-id: c849e36e1c4d1dce947076349df863fafe62c66d	2017-01-25 15:14:26 -08:00
Aapo Kyrola	95b3309a87	Gradient Input memory sharing using memonger blob sharing Summary: This diff brings us to roughly par with Torch on ResNet memory usage. On batch size 32, Resnet-50 took 7497MiB, after this 5010 MiB. This will thus allow us to handle 64 images / GPU, or 256 images / 4 GPUs. In addition, I added a special argument to DagNet that causes it to run only one thread for the first iteration. This is needed since there are allocations on the first iteration's backward pass due to gradient sharing, and this will cause NCCL to deadlock. The sharing of gradient buffers requires inferring which gradients can share memory (i.e that they are not used concurrently). Previous memonger code uses topological sort, but rbgirshick showed that it does not work with tree-like models. Thus, I wrote a new optimization algorithm based on DFS. It takes about 0.25 secs / GPU on resnet-50, so is clearly fast enough. Module data_parallel_model supports this feature natively. Reviewed By: prigoyal Differential Revision: D4363209 fbshipit-source-id: 73b11e7610438098bb11bff0af8075ab0cf2c0f1	2017-01-09 19:44:23 -08:00
Yangqing Jia	09bed67e4f	add untracked files	2016-07-21 11:26:41 -07:00

5 Commits