pytorch

mirror of https://github.com/zebrajr/pytorch.git synced 2025-12-07 00:21:07 +01:00

Author	SHA1	Message	Date
Catherine Lee	eea0733045	Reduce pytest blocklist (#96016 ) `TestCase = object` or variations of it get switched to `TestCase = NoTest`. unittest collects test based on subclassing unittest.TestCase, so setting TestCase = object removes it from unittest test collection. pytest collects based on name (https://docs.pytest.org/en/7.1.x/reference/reference.html#confval-python_classes) but can be told to ignore a class (bottom of https://docs.pytest.org/en/7.1.x/example/pythoncollection.html#changing-naming-conventions) Pull Request resolved: https://github.com/pytorch/pytorch/pull/96016 Approved by: https://github.com/ZainRizvi, https://github.com/huydhn	2023-03-07 18:30:27 +00:00
Eli Uriegas	b38b39c441	Revert "memory viz: Add colors for categories and a legend (#90587 )" This reverts commit `ee43842505`.	2023-03-06 11:38:58 -08:00
Zachary DeVito	ee43842505	memory viz: Add colors for categories and a legend (#90587 ) Adds a category legend to memory trace plots that colors allocations by their role (activation, parameter, gradient, etc.) as captured by kineto. Differential Revision: [D43757381](https://our.internmc.facebook.com/intern/diff/D43757381) Pull Request resolved: https://github.com/pytorch/pytorch/pull/90587 Approved by: https://github.com/aaronenyeshi	2023-03-03 20:42:22 +00:00
Mark Saroufim	9f707f164e	Add more GPU metric instrumentation (#91717 ) Fixes https://github.com/pytorch/serve/issues/1937 A fairly common query I see folks running while using pytorch is `nvidia-smi --format=csv,noheader,nounits --query-gpu=utilization.gpu,utilization.memory,memory.total,memory.used,temperature.gpu,power.draw,clocks.current.sm,clocks.current.memory -l 10` Existing metrics we have * For kernel utilization`torch.cuda.utilization()` * For memory utilization we have them under `torch.cuda.memory` the memory allocated with `torch.cuda.memory.memory_allocated()` * For total available memory we have `torch.cuda.get_device_properties(0).total_memory` Which means the only metrics we're missing are * Temperature: now in `torch.cuda.temperature()` * Power draw: now in `torch.cuda.power()` * Clock speed: now in `torch.cuda.clock_speed()` With some important details on each * Clock speed settings: I picked the SM clock domain which is documented here https://docs.nvidia.com/deploy/nvml-api/group__nvmlDeviceEnumvs.html#group__nvmlDeviceEnumvs_1g805c0647be9996589fc5e3f6ff680c64 * Temperature: I use `pynvml.nvmlDeviceGetTemperature(handle, 0)` where 0 refers to the GPU die temperature Pull Request resolved: https://github.com/pytorch/pytorch/pull/91717 Approved by: https://github.com/ngimel	2023-02-24 00:38:03 +00:00
Pearu Peterson	cece63f197	Add warn-once deprecation warning to legacy sparse constructors (#94850 ) Addresses https://github.com/pytorch/pytorch/issues/68323#issuecomment-1425174341 Pull Request resolved: https://github.com/pytorch/pytorch/pull/94850 Approved by: https://github.com/amjames, https://github.com/cpuhrsch	2023-02-23 15:05:12 +00:00
puririshi98	8aa34602f7	Jetson Update for CI Redo (#94549 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/94549 Approved by: https://github.com/ezyang, https://github.com/malfet	2023-02-21 17:13:38 +00:00
dllehr-amd	98012e4a59	[ROCm] hipGraph support for pytorch mainline (#88202 ) With the release of ROCm 5.3 hip now supports a hipGraph implementation. All necessary backend work and hipification is done to support the same functionality as cudaGraph. Unit tests are modified to support a new TEST_GRAPH feature which allows us to create a single check for graph support instead of attempted to gather the CUDA level in annotations for every graph test Pull Request resolved: https://github.com/pytorch/pytorch/pull/88202 Approved by: https://github.com/jithunnair-amd, https://github.com/pruthvistony, https://github.com/malfet	2023-02-14 22:18:56 +00:00
Xuehai Pan	b005ec62b9	[BE] Remove dependency on `six` and `future` (#94709 ) Remove the Python 2 and 3 compatibility library [six](https://pypi.org/project/six) and [future](https://pypi.org/project/future) and `torch._six`. We only support Python 3.8+ now. It's time to retire them. Pull Request resolved: https://github.com/pytorch/pytorch/pull/94709 Approved by: https://github.com/malfet, https://github.com/Skylion007	2023-02-14 09:14:14 +00:00
Xuehai Pan	046e88a291	[BE] [3/3] Rewrite `super()` calls in test (#94592 ) Rewrite Python built-in class `super()` calls. Only non-semantic changes should be applied. - #94587 - #94588 - #94592 Also, methods with only a `super()` call are removed: ```diff class MyModule(nn.Module): - def __init__(self): - super().__init__() - def forward(self, ...): ... ``` Some cases that change the semantics should be kept unchanged. E.g.: `f152a79be9/caffe2/python/net_printer.py (L184-L190)` `f152a79be9/test/test_jit_fuser_te.py (L2628-L2635)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/94592 Approved by: https://github.com/ezyang, https://github.com/seemethere	2023-02-12 22:20:53 +00:00
c-odrin	54b7c7d5e9	Added requested_bytes to CUDA Caching Allocator Stats (#88575 ) Summary: The caching allocator can be configured to round memory allocations in order to reduce fragmentation. Sometimes however, the overhead from rounding can be higher than the fragmentation it helps reduce. We have added a new stat to CUDA caching allocator stats to help track if rounding is adding too much overhead and help tune the roundup_power2_divisions flag: - "requested_bytes.{current,peak,allocated,freed}": memory requested by client code, compare this with allocated_bytes to check if allocation rounding adds too much overhead Test Plan: Added test case in caffe2/test/test_cuda.py Differential Revision: D40810674 Pull Request resolved: https://github.com/pytorch/pytorch/pull/88575 Approved by: https://github.com/zdevito	2023-02-09 21:37:25 +00:00
Masaki Kozuki	6ba041fcae	Look up `group["capturable"]`, not `defaults["capturable"]` in Adam(W) (#94149 ) We could set different values in each `param_group` when calling dunder init of `torch.optim` optimizers as in e.g. https://github.com/pytorch/pytorch/issues/89987. So check whether or not `capturable` is `True` among all the `param_group`s. Pull Request resolved: https://github.com/pytorch/pytorch/pull/94149 Approved by: https://github.com/albanD	2023-02-07 00:24:35 +00:00
Masaki Kozuki	4207d3c330	`FusedAdam(W)` should take `OptState` into account before unscaling grads (#94060 ) the optimizers have to consult `OptState` before unscaling gradients because we could call `GradScaler.unscale_` explicitly to for e.g. `clip_grad_norm_` as mentioned in `e52786f3d1/torch/cuda/amp/grad_scaler.py (L235-L266)` and https://pytorch.org/docs/stable/notes/amp_examples.html#working-with-unscaled-gradients Related #90752 Pull Request resolved: https://github.com/pytorch/pytorch/pull/94060 Approved by: https://github.com/albanD	2023-02-04 05:20:13 +00:00
Masaki Kozuki	a23ed38f9a	[mta][foreach] Implement fused adamw (#88015 ) related: https://github.com/pytorch/pytorch/issues/68041, https://github.com/pytorch/pytorch/issues/71274, https://github.com/pytorch/pytorch/issues/80167 possibly related to https://github.com/pytorch/pytorch/issues/80595#issuecomment-1178519436 Pull Request resolved: https://github.com/pytorch/pytorch/pull/88015 Approved by: https://github.com/albanD, https://github.com/ngimel	2023-02-01 19:32:29 +00:00
albanD	d8aa68c683	make sure that our error handling runs with the GIL enabled (#92848 ) Fixes https://github.com/pytorch/pytorch/issues/92684 I checked the other use case of this API and they never release the GIL Pull Request resolved: https://github.com/pytorch/pytorch/pull/92848 Approved by: https://github.com/ngimel	2023-01-24 09:30:42 +00:00
eqy	fb38b9ff2a	[cuBLAS][TF32] Fix TF32 get/set test when `TORCH_ALLOW_TF32_CUBLAS_OVERRIDE` is set (#92052 ) Follow up of #85859 to fix the test for when the environment variable is set. CC @xwang233 @ptrblck Pull Request resolved: https://github.com/pytorch/pytorch/pull/92052 Approved by: https://github.com/ngimel	2023-01-12 05:36:06 +00:00
eqy	97ff20d722	[cuBLAS] (re-open) Fix default cuBLAS workspace size and parsing for multiple workspaces (#91564 ) re-open of #89027 Pull Request resolved: https://github.com/pytorch/pytorch/pull/91564 Approved by: https://github.com/ngimel	2023-01-03 23:48:15 +00:00
PyTorch MergeBot	39d49dbe45	Revert "[cuBLAS] Fix default cuBLAS workspace size and parsing for multiple workspaces (#89027 )" This reverts commit `b407d98dbe`. Reverted https://github.com/pytorch/pytorch/pull/89027 on behalf of https://github.com/kit1980 due to Fails test_cublas_workspace_explicit_allocation on ROCm	2022-12-31 23:04:57 +00:00
eqy	b407d98dbe	[cuBLAS] Fix default cuBLAS workspace size and parsing for multiple workspaces (#89027 ) Follow-up of #86167 ; The number of pools was mistakenly ignored and the default workspace size appears to be too small to match selected cuBLAS kernels before the explicit allocation change. CC @ptrblck @ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/89027 Approved by: https://github.com/ngimel	2022-12-31 06:58:04 +00:00
lezcano	484dd40022	Implement PReLU in a compositional way (#91238 ) The PReLU implementation was all over the place. This lead to a number of bugs like https://github.com/pytorch/pytorch/issues/68760. We fix it by: - Keeping the weird broadcasting logic it has as a CompositeImplicit kernel that calls into a second kernel - This second kernel is just a good-ol' pointwise kernel. - We implement the derivative for the pointwise kernel via TI as well for speed. - We implement the second derivative for the pointwise kernel and the forward AD derivatives compositionally This fixes a number of issues: - We don't perform copies any more when the inputs are not contiguous - The derivatives are now correct - We fix vmap and many other functorch-related issues. - CPU and CUDA now share the relevant broadcasting logic - The implementation is about 1/3 the length. Fixes https://github.com/pytorch/pytorch/issues/68760 Fixes https://github.com/pytorch/pytorch/issues/89895 Pull Request resolved: https://github.com/pytorch/pytorch/pull/91238 Approved by: https://github.com/kshitij12345, https://github.com/jbschlosser, https://github.com/albanD	2022-12-30 10:42:30 +00:00
Eddie Yan	8b617f813d	[cuBLAS] Add an option to disable reduced precision reductions for BF16 GEMM (#89172 ) Essentially the same change as #67946, except that the default is to disallow reduced precision reductions in `BFloat16` GEMMs (for now). If performance is severely regressed, we can change the default, but this option appears to be necessary to pass some `addmm` `BFloat16` tests on H100. CC @ptrblck @ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/89172 Approved by: https://github.com/ngimel	2022-12-21 18:58:28 +00:00
Eddie Yan	dabf515c18	[cuDNN][cuDNN V8 API] (re-re-re-open) cuDNN V8 API on by default (#91117 ) Re-opening following #91025 CC @ptrblck @ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/91117 Approved by: https://github.com/ngimel	2022-12-20 18:52:29 +00:00
PyTorch MergeBot	ba7aeac37b	Revert "[cuDNN][cuDNN V8 API] (re-re-open) cuDNN V8 API on by default (#89022 )" This reverts commit `eecd621f06`. Reverted https://github.com/pytorch/pytorch/pull/89022 on behalf of https://github.com/ngimel due to breaks some convolution configurations #91025	2022-12-16 23:06:35 +00:00
Rich Zhu	4372dbb89f	use pytree to allow any input format for cuda graph (#90941 ) Summary: 1. use pytree to allow any input format for make_graphed_callables 2. add allow_unused_input argument for make_graphed_callables Test Plan: buck2 test mode/dev-nosan //caffe2/test:cuda -- --print-passing-details Differential Revision: D42077976 Pull Request resolved: https://github.com/pytorch/pytorch/pull/90941 Approved by: https://github.com/ngimel	2022-12-16 03:01:47 +00:00
Eddie Yan	eecd621f06	[cuDNN][cuDNN V8 API] (re-re-open) cuDNN V8 API on by default (#89022 ) Testing V8 on by default again after fixes have been merged for e.g., https://github.com/pytorch/torchdynamo/issues/1833 One new failure that seems to be surfaced with V8 on appears in halonext + amp ``` RuntimeError: Internal Triton PTX codegen error: Segmentation fault (core dumped) ``` But I'm not sure if this points to a V8 issue or a Triton issue CC @ngimel @ptrblck Current dynamo benchmarks on A100: v7 vs. v8 \|dev \|name \|batch_size\|abs_latency_v7\|abs_latency_v8\| \|----\|-------------------------------\|----------\|--------------\|--------------\| \|cuda\|adv_inception_v3 \|128 \|166.0240 \|165.5798 \| \|cuda\|beit_base_patch16_224 \|64 \|123.5912 \|123.0797 \| \|cuda\|botnet26t_256 \|128 \|107.7343 \|107.5948 \| \|cuda\|cait_m36_384 \|4 \|184.5038 \|184.0271 \| \|cuda\|coat_lite_mini \|128 \|142.3061 \|140.5814 \| \|cuda\|convit_base \|64 \|165.2499 \|161.0743 \| \|cuda\|convmixer_768_32 \|32 \|325.6984 \|325.7094 \| \|cuda\|convnext_base \|64 \|237.4632 \|238.0142 \| \|cuda\|crossvit_9_240 \|128 \|72.2980 \|72.4367 \| \|cuda\|cspdarknet53 \|64 \|96.6862 \|96.8308 \| \|cuda\|deit_base_distilled_patch16_224\|64 \|117.6045 \|117.9616 \| \|cuda\|dla102 \|128 \|182.3073 \|182.2304 \| \|cuda\|dm_nfnet_f0 \|128 \|133.6011 \|133.6298 \| \|cuda\|dpn107 \|32 \|148.5080 \|148.5885 \| \|cuda\|eca_botnext26ts_256 \|128 \|113.8676 \|113.1514 \| \|cuda\|eca_halonext26ts \|128 \|119.2242 \|119.1845 \| \|cuda\|ese_vovnet19b_dw \|128 \|80.0217 \|79.9438 \| \|cuda\|fbnetc_100 \|128 \|91.4548 \|91.4009 \| \|cuda\|fbnetv3_b \|128 \|115.4496 \|115.5058 \| \|cuda\|gernet_l \|128 \|114.8365 \|114.7870 \| \|cuda\|ghostnet_100 \|128 \|58.5766 \|58.5766 \| \|cuda\|gluon_inception_v3 \|128 \|165.5222 \|165.7167 \| \|cuda\|gluon_xception65 \|32 \|165.8779 \|165.7818 \| \|cuda\|gmixer_24_224 \|128 \|116.3611 \|113.4925 \| \|cuda\|gmlp_s16_224 \|128 \|121.2607 \|121.2534 \| \|cuda\|hrnet_w18 \|128 \|246.5706 \|246.7599 \| \|cuda\|inception_v3 \|128 \|166.1096 \|166.2034 \| \|cuda\|jx_nest_base \|32 \|93.6064 \|93.4088 \| \|cuda\|lcnet_050 \|128 \|21.4156 \|21.4207 \| \|cuda\|levit_128 \|128 \|27.2901 \|27.2543 \| \|cuda\|mixer_b16_224 \|128 \|157.8992 \|158.2878 \| \|cuda\|mixnet_l \|128 \|197.3443 \|197.2125 \| \|cuda\|mnasnet_100 \|128 \|71.4604 \|71.2997 \| \|cuda\|mobilenetv2_100 \|128 \|67.6080 \|67.7515 \| \|cuda\|mobilenetv3_large_100 \|128 \|57.7224 \|57.6591 \| \|cuda\|mobilevit_s \|64 \|93.0372 \|93.0530 \| \|cuda\|nfnet_l0 \|128 \|113.1664 \|113.2853 \| \|cuda\|pit_b_224 \|64 \|133.3333 \|133.4153 \| \|cuda\|pnasnet5large \|16 \|238.9545 \|238.8122 \| \|cuda\|poolformer_m36 \|64 \|144.2353 \|144.2375 \| \|cuda\|regnety_002 \|128 \|32.8534 \|32.9069 \| \|cuda\|repvgg_a2 \|128 \|102.4150 \|102.3827 \| \|cuda\|res2net101_26w_4s \|64 \|120.8127 \|120.8322 \| \|cuda\|res2net50_14w_8s \|128 \|149.7052 \|149.8969 \| \|cuda\|res2next50 \|128 \|153.7439 \|153.8215 \| \|cuda\|resmlp_12_224 \|128 \|89.1918 \|86.9226 \| \|cuda\|resnest101e \|64 \|159.4706 \|159.3133 \| \|cuda\|rexnet_100 \|128 \|88.0032 \|88.0397 \| \|cuda\|sebotnet33ts_256 \|64 \|80.4635 \|80.0120 \| \|cuda\|selecsls42b \|128 \|70.4430 \|70.3663 \| \|cuda\|spnasnet_100 \|128 \|78.0537 \|78.1991 \| \|cuda\|swin_base_patch4_window7_224 \|64 \|212.9073 \|213.0824 \| \|cuda\|swsl_resnext101_32x16d \|32 \|193.0229 \|193.0404 \| \|cuda\|tf_efficientnet_b0 \|128 \|97.1316 \|97.0410 \| \|cuda\|tf_mixnet_l \|128 \|203.4956 \|203.5340 \| \|cuda\|tinynet_a \|128 \|82.4038 \|82.8733 \| \|cuda\|tnt_s_patch16_224 \|128 \|284.8576 \|284.8867 \| \|cuda\|twins_pcpvt_base \|64 \|118.3893 \|119.2329 \| \|cuda\|visformer_small \|128 \|126.0533 \|126.0390 \| \|cuda\|vit_base_patch16_224 \|64 \|118.2873 \|118.0573 \| \|cuda\|volo_d1_224 \|64 \|108.7764 \|108.2063 \| \|cuda\|xcit_large_24_p8_224 \|5 \|100.4656 \|100.5209 \| v7 vs. v8 amp \|dev \|name \|batch_size\|abs_latency_v7\|abs_latency_v8\| \|----\|-------------------------------\|----------\|--------------\|--------------\| \|cuda\|adv_inception_v3 \|128 \|104.9729 \|105.1237 \| \|cuda\|beit_base_patch16_224 \|64 \|75.4330 \|75.2039 \| \|cuda\|botnet26t_256 \|128 \|74.5149 \|74.8071 \| \|cuda\|cait_m36_384 \|4 \|110.9788 \|111.5170 \| \|cuda\|coat_lite_mini \|128 \|62.3618 \|64.4965 \| \|cuda\|convit_base \|64 \|116.4054 \|117.9129 \| \|cuda\|convmixer_768_32 \|32 \|264.4401 \|264.4491 \| \|cuda\|convnext_base \|64 \|182.9009 \|179.2136 \| \|cuda\|crossvit_9_240 \|128 \|48.8586 \|48.8359 \| \|cuda\|cspdarknet53 \|64 \|80.0245 \|80.0160 \| \|cuda\|deit_base_distilled_patch16_224\|64 \|66.5921 \|66.7448 \| \|cuda\|dla102 \|128 \|116.7780 \|117.1683 \| \|cuda\|dm_nfnet_f0 \|128 \|78.9322 \|79.1135 \| \|cuda\|dpn107 \|32 \|85.5206 \|85.7514 \| \|cuda\|eca_botnext26ts_256 \|128 \|76.3672 \|77.0050 \| \|cuda\|eca_halonext26ts \|128 \|86.2458 \| \| \|cuda\|ese_vovnet19b_dw \|128 \|43.2943 \|43.3379 \| \|cuda\|fbnetc_100 \|128 \|54.8479 \|54.9251 \| \|cuda\|fbnetv3_b \|128 \|70.7504 \|71.0188 \| \|cuda\|gernet_l \|128 \|66.1607 \|66.0379 \| \|cuda\|ghostnet_100 \|128 \|43.8882 \|43.9336 \| \|cuda\|gluon_inception_v3 \|128 \|104.9297 \|105.0204 \| \|cuda\|gluon_xception65 \|32 \|85.7118 \|85.8370 \| \|cuda\|gmixer_24_224 \|128 \|75.1214 \|76.1170 \| \|cuda\|gmlp_s16_224 \|128 \|76.4207 \|76.6641 \| \|cuda\|hrnet_w18 \|128 \|186.1326 \|186.2435 \| \|cuda\|inception_v3 \|128 \|105.0561 \|105.0783 \| \|cuda\|jx_nest_base \|32 \|65.3066 \|65.3245 \| \|cuda\|lcnet_050 \|128 \|14.7991 \|14.8687 \| \|cuda\|levit_128 \|128 \|19.2893 \|19.4772 \| \|cuda\|mixer_b16_224 \|128 \|93.9826 \|94.2056 \| \|cuda\|mixnet_l \|128 \|147.1245 \|147.0435 \| \|cuda\|mnasnet_100 \|128 \|39.1781 \|39.2565 \| \|cuda\|mobilenetv2_100 \|128 \|42.3704 \|42.3114 \| \|cuda\|mobilenetv3_large_100 \|128 \|37.2946 \|37.2816 \| \|cuda\|mobilevit_s \|64 \|55.8930 \|55.8934 \| \|cuda\|nfnet_l0 \|128 \|64.0448 \|64.4438 \| \|cuda\|pit_b_224 \|64 \|80.6342 \|80.2933 \| \|cuda\|pnasnet5large \|16 \|154.9611 \|154.8654 \| \|cuda\|poolformer_m36 \|64 \|101.7489 \|101.8138 \| \|cuda\|regnety_002 \|128 \|27.0939 \|27.0309 \| \|cuda\|repvgg_a2 \|128 \|60.9651 \|61.2533 \| \|cuda\|res2net101_26w_4s \|64 \|77.3291 \|77.4739 \| \|cuda\|res2net50_14w_8s \|128 \|93.6572 \|93.7221 \| \|cuda\|res2next50 \|128 \|112.4975 \|112.3248 \| \|cuda\|resmlp_12_224 \|128 \|59.5422 \|60.7644 \| \|cuda\|resnest101e \|64 \|97.9894 \|98.3358 \| \|cuda\|rexnet_100 \|128 \|55.2218 \|55.0718 \| \|cuda\|sebotnet33ts_256 \|64 \|60.4880 \|60.8113 \| \|cuda\|selecsls42b \|128 \|41.4294 \|41.5341 \| \|cuda\|spnasnet_100 \|128 \|45.0037 \|45.0304 \| \|cuda\|swin_base_patch4_window7_224 \|64 \|98.2561 \|98.6925 \| \|cuda\|swsl_resnext101_32x16d \|32 \|100.6179 \|100.9195 \| \|cuda\|tf_efficientnet_b0 \|128 \|56.5344 \|56.4591 \| \|cuda\|tf_mixnet_l \|128 \|153.0318 \|152.9367 \| \|cuda\|tinynet_a \|128 \|54.1307 \|53.9298 \| \|cuda\|tnt_s_patch16_224 \|128 \|142.4801 \|142.6589 \| \|cuda\|twins_pcpvt_base \|64 \|67.9027 \|67.8325 \| \|cuda\|visformer_small \|128 \|72.5589 \|72.9427 \| \|cuda\|vit_base_patch16_224 \|64 \|71.4885 \|71.7342 \| \|cuda\|volo_d1_224 \|64 \|69.3539 \|69.5910 \| \|cuda\|xcit_large_24_p8_224 \|5 \|59.9000 \|59.9699 \| v7 vs. v8 float16 \|dev \|name \|batch_size\|abs_latency\|abs_latency\| \|----\|-------------------------------\|----------\|-----------\|-----------\| \|cuda\|adv_inception_v3 \|128 \|104.2544 \|104.2677 \| \|cuda\|beit_base_patch16_224 \|64 \|85.3601 \|85.3786 \| \|cuda\|botnet26t_256 \|128 \|72.1476 \|71.8277 \| \|cuda\|cait_m36_384 \|4 \|108.3075 \|108.5941 \| \|cuda\|coat_lite_mini \|128 \|61.2382 \|61.6049 \| \|cuda\|convmixer_768_32 \|32 \|263.3818 \|263.3598 \| \|cuda\|convnext_base \|64 \|172.6821 \|173.8520 \| \|cuda\|crossvit_9_240 \|128 \|44.6321 \|44.6340 \| \|cuda\|cspdarknet53 \|64 \|79.3165 \|79.2964 \| \|cuda\|deit_base_distilled_patch16_224\|64 \|61.9816 \|62.2109 \| \|cuda\|dla102 \|128 \|115.7403 \|115.9928 \| \|cuda\|dm_nfnet_f0 \|128 \|77.5434 \|77.7440 \| \|cuda\|dpn107 \|32 \|83.6489 \|83.5605 \| \|cuda\|eca_botnext26ts_256 \|128 \|73.9953 \|74.1031 \| \|cuda\|eca_halonext26ts \|128 \|81.7951 \|81.7103 \| \|cuda\|ese_vovnet19b_dw \|128 \|42.9618 \|42.8853 \| \|cuda\|fbnetc_100 \|128 \|54.3590 \|54.3575 \| \|cuda\|fbnetv3_b \|128 \|69.7977 \|70.1696 \| \|cuda\|gernet_l \|128 \|64.8684 \|65.1726 \| \|cuda\|ghostnet_100 \|128 \|43.2054 \|43.1319 \| \|cuda\|gluon_inception_v3 \|128 \|104.1988 \|104.3030 \| \|cuda\|gluon_xception65 \|32 \|84.2245 \|84.5085 \| \|cuda\|gmixer_24_224 \|128 \|82.0418 \|82.7252 \| \|cuda\|gmlp_s16_224 \|128 \|75.4792 \|75.8374 \| \|cuda\|hrnet_w18 \|128 \|184.1450 \|184.1848 \| \|cuda\|inception_v3 \|128 \|104.1203 \|104.2536 \| \|cuda\|jx_nest_base \|32 \|58.2386 \|58.4901 \| \|cuda\|lcnet_050 \|128 \|14.6409 \|14.5616 \| \|cuda\|levit_128 \|128 \|22.3875 \|22.4680 \| \|cuda\|mixer_b16_224 \|128 \|98.9534 \|98.4730 \| \|cuda\|mixnet_l \|128 \|146.1623 \|146.1947 \| \|cuda\|mnasnet_100 \|128 \|38.9208 \|39.3463 \| \|cuda\|mobilenetv2_100 \|128 \|41.8946 \|41.9847 \| \|cuda\|mobilenetv3_large_100 \|128 \|36.7810 \|36.8264 \| \|cuda\|mobilevit_s \|64 \|55.3211 \|55.3186 \| \|cuda\|nfnet_l0 \|128 \|63.1302 \|63.5544 \| \|cuda\|pit_b_224 \|64 \|73.8752 \|73.4602 \| \|cuda\|pnasnet5large \|16 \|151.6806 \|151.6111 \| \|cuda\|poolformer_m36 \|64 \|86.8341 \|86.8021 \| \|cuda\|regnety_002 \|128 \|26.6798 \|26.5295 \| \|cuda\|repvgg_a2 \|128 \|61.6652 \|62.1482 \| \|cuda\|res2net101_26w_4s \|64 \|75.8037 \|75.7739 \| \|cuda\|res2net50_14w_8s \|128 \|92.6362 \|92.4338 \| \|cuda\|res2next50 \|128 \|111.5371 \|111.5832 \| \|cuda\|resmlp_12_224 \|128 \|58.2349 \|57.9807 \| \|cuda\|resnest101e \|64 \|96.1114 \|96.2742 \| \|cuda\|rexnet_100 \|128 \|54.8138 \|54.7643 \| \|cuda\|sebotnet33ts_256 \|64 \|53.1524 \|53.3823 \| \|cuda\|selecsls42b \|128 \|40.6070 \|40.7104 \| \|cuda\|spnasnet_100 \|128 \|44.5732 \|44.4318 \| \|cuda\|swin_base_patch4_window7_224 \|64 \|98.6447 \|98.8445 \| \|cuda\|swsl_resnext101_32x16d \|32 \|97.0195 \|97.2968 \| \|cuda\|tf_efficientnet_b0 \|128 \|56.0640 \|56.0278 \| \|cuda\|tf_mixnet_l \|128 \|152.0958 \|152.0874 \| \|cuda\|tinynet_a \|128 \|53.3694 \|53.3762 \| \|cuda\|tnt_s_patch16_224 \|128 \|130.2981 \|130.3726 \| \|cuda\|twins_pcpvt_base \|64 \|62.5459 \|62.6416 \| \|cuda\|visformer_small \|128 \|68.8502 \|69.1756 \| \|cuda\|vit_base_patch16_224 \|64 \|65.8587 \|66.0285 \| \|cuda\|volo_d1_224 \|64 \|64.5348 \|64.6057 \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/89022 Approved by: https://github.com/ngimel	2022-12-15 03:24:44 +00:00
PyTorch MergeBot	cba96366a2	Revert "remove torch.equal usages (#89527 )" This reverts commit `4095ef8b80`. Reverted https://github.com/pytorch/pytorch/pull/89527 on behalf of https://github.com/clee2000 due to broke periodic multigpu tests `4095ef8b80` https://github.com/pytorch/pytorch/actions/runs/3592806602/jobs/6049368502	2022-12-02 21:36:13 +00:00
Philip Meier	4095ef8b80	remove torch.equal usages (#89527 ) Preparation for the next PR in this stack: #89559. I replaced - `self.assertTrue(torch.equal(...))` with `self.assertEqual(..., rtol=0, atol=0, exact_device=True)`, - the same for `self.assertFalse(...)` with `self.assertNotEqual(...)`, and - `assert torch.equal(...)` with `torch.testing.assert_close(..., rtol=0, atol=0)` (note that we don't need to set `check_device=True` here since that is the default). There were a few instances where the result of `torch.equal` is used directly. In that cases I've replaced with `(... == ...).all().item()` while sometimes also dropping the `.item()` depending on the context. Pull Request resolved: https://github.com/pytorch/pytorch/pull/89527 Approved by: https://github.com/mruberry	2022-12-01 11:22:52 +00:00
Aidyn-A	0057be3361	[CUDA graphs] Add warning if captured graph is empty (#88754 ) Fixes #87894 This PR adds a warning if captured graph is empty (consists of zero nodes). The example snippet where would it be useful: ```python import torch x = torch.randn(10) z = torch.zeros(10) g = torch.cuda.CUDAGraph() with torch.cuda.graph(g): z = x * x # Warn user ``` and in #87894 Pull Request resolved: https://github.com/pytorch/pytorch/pull/88754 Approved by: https://github.com/ezyang	2022-11-28 23:20:19 +00:00
Nikita Shulga	da2afcb1e0	Add test for out-of-bounds Tensor access on GPU (#39211 ) Since CUDA context can not recover safely from on-device assert, use `torch.multiprocessing.spawn` to execute a method in another context and verify that it raises unrecoverable error. As those types of tests are pretty slow (6 seconds on powerful linux box with one GPU) run it only in the slow shard. Closes https://github.com/pytorch/pytorch/issues/38944 Pull Request resolved: https://github.com/pytorch/pytorch/pull/39211 Approved by: https://github.com/ezyang	2022-11-15 21:06:02 +00:00
PyTorch MergeBot	d98a884b33	Revert "[cuDNN] (re-open) Enable cuDNN Frontend v8 API by Default (#87669 )" This reverts commit `3c6bddc3f6`. Reverted https://github.com/pytorch/pytorch/pull/87669 on behalf of https://github.com/eqy due to investigating convnext benchmark regressions	2022-11-08 19:04:25 +00:00
Kurt Mohler	ee28b865ee	Deprecate TypedStorage, its derived classes, and all of their public methods (#85303 ) Part of #85302 Pull Request resolved: https://github.com/pytorch/pytorch/pull/85303 Approved by: https://github.com/ezyang	2022-11-08 18:11:01 +00:00
Codrin Popa	5b767d404e	Modified roundup_power2_divisions to specify the number of divisions for each power of two interval (#87290 ) Summary: Improved roundup_power2_divisions knob so it allows better control of rouding in the PyTorch CUDA Caching Allocator. This new version allows setting the number of divisions per power of two interval starting from 1MB and ending at 64GB and above. An example use case is when rouding is desirable for small allocations but there are also very large allocations which are persistent, thus would not benefit from rounding and take up extra space. Test Plan: Tested locally Differential Revision: D40103909 Pull Request resolved: https://github.com/pytorch/pytorch/pull/87290 Approved by: https://github.com/zdevito	2022-11-04 19:31:16 +00:00
eqy	3c6bddc3f6	[cuDNN] (re-open) Enable cuDNN Frontend v8 API by Default (#87669 ) #58414 Has a small tweak to a test that was breaking on A10 (CC @malfet). CC @ptrblck @ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/87669 Approved by: https://github.com/ngimel	2022-11-02 01:36:37 +00:00
Masaki Kozuki	bc03aa6013	Store `autocast_gpu_dtype` in `custom_fwd` and `custom_bwd` for BFloat16 autocast (#88029 ) As per #87979, `custom_bwd` seems to forcefully use `torch.float16` for `torch.autograd.Function.backward` regardless of the `dtype` used in the forward. Changes: - store the `dtype` in `args[0]` - update tests to confirm the dtype of intermediate result tensors that are outputs of autocast compatible `torch` functions cc @ptrblck @ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/88029 Approved by: https://github.com/ngimel	2022-10-31 22:45:26 +00:00
Zachary DeVito	00c91f4446	[allocator] disable tests that don't work for cudaMallocAsyncAllocator (#87250 ) Two tests were failing locally for me and don't appear to be run in our CI. Disabling them so we can otherwise refactor the allocators. Pull Request resolved: https://github.com/pytorch/pytorch/pull/87250 Approved by: https://github.com/wconstab	2022-10-19 18:29:35 +00:00
PyTorch MergeBot	746500d58d	Revert "[cuDNN] Enable cuDNN Frontend v8 API by Default (#84948 )" This reverts commit `427e0a6b4e`. Reverted https://github.com/pytorch/pytorch/pull/84948 on behalf of https://github.com/malfet due to Broke SM86 sanity	2022-10-14 14:25:51 +00:00
Eddie Yan	427e0a6b4e	[cuDNN] Enable cuDNN Frontend v8 API by Default (#84948 ) #58414 Opening this PR for testing for now to check CI status. 🤞 CC @ptrblck @ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/84948 Approved by: https://github.com/ngimel	2022-10-13 17:26:36 +00:00
Eddie Yan	25725fd624	(Re-open) Adds cudaMallocAsync as an alternative backend for the CUDA allocator (#82682 ) Rebased version of @mcarilli 's cudaMallocAsync #65365 for continued testing Pull Request resolved: https://github.com/pytorch/pytorch/pull/82682 Approved by: https://github.com/ngimel	2022-10-12 03:44:21 +00:00
eqy	352d926482	[CUBLAS][CUDA GRAPHS] (re-re-re-re-open of #83461 ) Explicitly set the workspace for cuBLAS handles (#86645 ) re-opening (again) in hopes of working around failed/stuck CLA check CC @ptrblck @ngimel @huydhn Pull Request resolved: https://github.com/pytorch/pytorch/pull/86645 Approved by: https://github.com/zdevito	2022-10-11 16:03:49 +00:00
Zachary DeVito	91b1bae1df	Caching allocator tracing (#86241 ) We currently can take snapshots of the state of the allocated cuda memory, but we do not have a way to correlate these snapshots with the actions the allocator that were taken between snapshots. This PR adds a simple fixed-sized buffer that records the major actions that the allocator takes (ALLOC, FREE, SEGMENT_ALLOC, SEGMENT_FREE, OOM, SNAPSHOT) and includes these with the snapshot information. Capturing period snapshots with a big enough trace buffer makes it possible to see how the allocator state changes over time. We plan to use this functionality to guide how settings in the allocator can be adjusted and eventually have a more robust overall algorithm. As a component of this functionality, we also add the ability to get a callback when the allocator will throw an OOM, primarily so that snapshots can be taken immediately to see why the program ran out of memory (most programs have some C++ state that would free tensors before the OutOfMemory exception can be caught). This PR also updates the _memory_viz.py script to pretty-print the trace information and provide a better textual summary of snapshots distinguishing between internal and external fragmentation. Pull Request resolved: https://github.com/pytorch/pytorch/pull/86241 Approved by: https://github.com/ngimel	2022-10-07 23:19:54 +00:00
Edward Z. Yang	adf5919720	Add option to record C++ backtraces in _record_memory_history (#86145 ) I used this to debug https://github.com/pytorch/pytorch/issues/86136 so it is useful. The implementation is not so fast so it is not enabled by default. Signed-off-by: Edward Z. Yang <ezyang@fb.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/86145 Approved by: https://github.com/albanD, https://github.com/zdevito	2022-10-06 04:07:37 +00:00
Zachary DeVito	736adc0808	Memory snapshots from C++ (#86190 ) Sometimes the driving process want to save memory snapshots but isn't Python. Add a simple API to turn it on without python stack traces. It still saves to the same format for the vizualization and summary scripts, using the C++ Pickler. Pull Request resolved: https://github.com/pytorch/pytorch/pull/86190 Approved by: https://github.com/ezyang	2022-10-05 07:36:39 +00:00
PyTorch MergeBot	71eb04403c	Revert "[CUBLAS][CUDA GRAPHS] (re-re-open of #83461 ) Explicitly set the workspace for cuBLAS handles (#85447 )" This reverts commit `b04b2fa9aa`. Reverted https://github.com/pytorch/pytorch/pull/85447 on behalf of https://github.com/seemethere due to Caused a CUDA memory leak, detected by our performance benchmark suite	2022-09-30 20:53:41 +00:00
Masaki Kozuki	5f26df0345	resubmit: "resubmit: [mta] APEX style Fused Adam (#81705 ) (#85507 )" (#85739 ) Embarrassingly move the pow implementations around [ATen/native/cuda/PowKernel.cu#L21-L66](`849b08f14b/aten/src/ATen/native/cuda/PowKernel.cu (L21-L66)`) to a new header file and let FusedAdam use them to tame MSVC, hopefully. cc @ngimel @ptrblck Pull Request resolved: https://github.com/pytorch/pytorch/pull/85739 Approved by: https://github.com/ngimel	2022-09-29 16:58:59 +00:00
Eddie Yan	b04b2fa9aa	[CUBLAS][CUDA GRAPHS] (re-re-open of #83461 ) Explicitly set the workspace for cuBLAS handles (#85447 ) Now includes @dagitses 's optimizations and fixes for teardown CC @ngimel @ptrblck Pull Request resolved: https://github.com/pytorch/pytorch/pull/85447 Approved by: https://github.com/malfet	2022-09-28 16:04:58 +00:00
Andres Lugo-Reyes	5709c67f1f	[ROCm] Retry loop implemented to avoid transient memory leak errors (#82607 ) ### Description Added a retry loop to memory leak checker to avoid rare case in which ROCM reports a false positive memory leak. ### Issue Original issue observed as part of this ticket: https://github.com/pytorch/pytorch/issues/62533 ### Testing - Applied changes and built - python test/test_cuda.py - Ensure all tests pass Pull Request resolved: https://github.com/pytorch/pytorch/pull/82607 Approved by: https://github.com/malfet	2022-09-28 15:48:24 +00:00
PyTorch MergeBot	7167996346	Revert "resubmit: [mta] APEX style Fused Adam (#81705 ) (#85507 )" This reverts commit `4615d1bcfa`. Reverted https://github.com/pytorch/pytorch/pull/85507 on behalf of https://github.com/atalman due to Break internal windows builds	2022-09-27 16:59:35 +00:00
Masaki Kozuki	4615d1bcfa	resubmit: [mta] APEX style Fused Adam (#81705 ) (#85507 ) This PR implements an APEX style FusedAdam in PyTorch. This is different from the APEX one in that this is compatible with `torch.cuda.amp.GradScaler` by setting `_step_supports_amp_scaling` to `True` and unscales gradients inside its CUDA kernel. related: https://github.com/pytorch/pytorch/issues/68041, https://github.com/pytorch/pytorch/issues/71274, https://github.com/pytorch/pytorch/issues/80167 possibly related to https://github.com/pytorch/pytorch/issues/80595#issuecomment-1178519436 Pull Request resolved: https://github.com/pytorch/pytorch/pull/81705 Approved by: https://github.com/ngimel cc @ptrblck @ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/85507 Approved by: https://github.com/ngimel	2022-09-23 18:56:00 +00:00
PyTorch MergeBot	e505360eb8	Revert "[mta] APEX style Fused Adam (#81705 )" This reverts commit `7a6c4d0c50`. Reverted https://github.com/pytorch/pytorch/pull/81705 on behalf of https://github.com/dagitses due to broke internal builds, details to come	2022-09-22 19:37:29 +00:00
PyTorch MergeBot	0ac6311356	Revert "[CUBLAS][CUDA GRAPHS] (re-open of #83461 ) Explicitly set the workspace for cuBLAS handles (#85292 )" This reverts commit `4012e623e8`. Reverted https://github.com/pytorch/pytorch/pull/85292 on behalf of https://github.com/dagitses due to broke an internal test during shutdown. Re-submit with #85399 in stack	2022-09-21 17:57:49 +00:00
Masaki Kozuki	7a6c4d0c50	[mta] APEX style Fused Adam (#81705 ) This PR implements an APEX style FusedAdam in PyTorch. This is different from the APEX one in that this is compatible with `torch.cuda.amp.GradScaler` by setting `_step_supports_amp_scaling` to `True` and unscales gradients inside its CUDA kernel. related: https://github.com/pytorch/pytorch/issues/68041, https://github.com/pytorch/pytorch/issues/71274, https://github.com/pytorch/pytorch/issues/80167 possibly related to https://github.com/pytorch/pytorch/issues/80595#issuecomment-1178519436 cc @ptrblck @ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/81705 Approved by: https://github.com/ngimel	2022-09-20 17:18:33 +00:00
eqy	4012e623e8	[CUBLAS][CUDA GRAPHS] (re-open of #83461 ) Explicitly set the workspace for cuBLAS handles (#85292 ) re-open of #83461 with fix for 10.2 build CC @ngimel @malfet Pull Request resolved: https://github.com/pytorch/pytorch/pull/85292 Approved by: https://github.com/malfet	2022-09-20 16:31:54 +00:00
Hector Yuen	d23ce29761	allow changing the cuda allocator settings even after the process started (#84970 ) Summary: - expose a python call to set the allocator settings, it uses the same format as the value for PYTORCH_CUDA_ALLOCATOR - keep the implementation contained within the cpp file to avoid increasing build times, only expose a function to call the setting - make some of the Allocator Config methods public, now it looks more like a singleton Test Plan: added the unit test Differential Revision: D39487522 Pull Request resolved: https://github.com/pytorch/pytorch/pull/84970 Approved by: https://github.com/zdevito	2022-09-17 09:42:42 +00:00
PyTorch MergeBot	2711b9fa63	Revert "[CUBLAS][CUDA GRAPHS] Explicitly set the workspace for cuBLAS handles (#83461 )" This reverts commit `713d8b8552`. Reverted https://github.com/pytorch/pytorch/pull/83461 on behalf of https://github.com/malfet due to Broke CUDA-10.2 builds, see `713d8b8552`	2022-09-14 22:27:30 +00:00
Eddie Yan	713d8b8552	[CUBLAS][CUDA GRAPHS] Explicitly set the workspace for cuBLAS handles (#83461 ) We're seeing an issue where repeatedly capturing graphs incurs increasing memory usage as cuBLAS internally allocates a new workspace for each graph even when the same handle is being used: https://gist.github.com/tomconerlyanth/a20c04a4a46a0f6e9ce18f5280729b36 This PR works around the issue by intercepting the `CUBLAS_WORKSPACE_CONFIG` environment variable and allocating the workspace for the cuBLAS handle explicitly. CC @ptrblck Pull Request resolved: https://github.com/pytorch/pytorch/pull/83461 Approved by: https://github.com/ngimel	2022-09-14 21:56:48 +00:00
Aidyn-A	5271494ef2	[CUDA graphs] Fixes errors in RNG seed (#84967 ) Fixes #84614 Prior to this PR CUDAGraph did not store the RNG seed, that is why `torch.cuda.manual_seed(new_seed)` would only reset the offset but not update the seed at all keeping whatever value was used during graph capture. Pull Request resolved: https://github.com/pytorch/pytorch/pull/84967 Approved by: https://github.com/ngimel	2022-09-14 19:56:12 +00:00
jataylo	09bcc006e9	ROCm support for test_lazy_init (#84333 ) Added ROCm support for the test_lazy_init unit test by including a condition on TEST_WITH_ROCM to switch CUDA_VISIBLE_DEVICES with HIP_VISIBLE_DEVICES. This is needed because HIP_VISIBLE_DEVICES is set when running the single-GPU tests in CI: `a47bc96fb7/.jenkins/pytorch/test.sh (L38)`, but this test sets CUDA_VISIBLE_DEVICES, which takes lower precedence than HIP_VISIBLE_DEVICES on ROCm. Testing Logs (to show behavior difference) 12:40:41 Aug 30 11:40:41 CUDA_VISIBLE_DEVICES='0': 0 12:40:41 Aug 30 11:40:41 1 12:40:41 Aug 30 11:40:41 CUDA_VISIBLE_DEVICES='32': 32 12:40:41 Aug 30 11:40:41 1 12:40:41 Aug 30 11:40:41 HIP_VISIBLE_DEVICES='0': 0 12:40:41 Aug 30 11:40:41 1 12:40:41 Aug 30 11:40:41 HIP_VISIBLE_DEVICES='32': 32 12:40:41 Aug 30 11:40:41 0 Passing UT Aug 30 17:03:15 test_lazy_init (main.TestCuda) Aug 30 17:03:17 Validate that no CUDA calls are made during import torch call ... ok (2.471s) Pull Request resolved: https://github.com/pytorch/pytorch/pull/84333 Approved by: https://github.com/jithunnair-amd, https://github.com/malfet	2022-09-09 14:14:59 +00:00
Fabio Rocha	88b1cc885c	Removed tri[lu]* tests, superseeded by OpInfos (#84256 ) triu, tril, triu_indices and tril_indices had some tests in test_tensor_creation_ops.py and test_cuda.py that are redudant with the ones done by OpInfos for those ops. Pull Request resolved: https://github.com/pytorch/pytorch/pull/84256 Approved by: https://github.com/Lezcano, https://github.com/ngimel	2022-09-06 18:54:10 +00:00
Aidyn-A	ce1b727e77	Disable autocast cache in torch.cuda.make_graphed_callables (#84289 ) There there are conflicts between `torch.clear_autocast_cache()` and `cudaMallocAsync` from #82682. Moreover, the use of autocast caching is not reasonable during training which is the main target of `make_graphed_callables`. cc @eqy @ptrblck Pull Request resolved: https://github.com/pytorch/pytorch/pull/84289 Approved by: https://github.com/ngimel	2022-09-01 21:34:51 +00:00
Pruthvi Madugundu	8473e69684	[ROCm] Fixes the kernel asserts API declaration mismatch error (#81790 ) This problem updates the the PR [#73040](https://github.com/pytorch/pytorch/pull/73040) The compilation error in pyTorch with ROCm is successful with these changes when `NDEBUG` is enabled. Solution: For HIP we keep `__device__ __assert_fail()` and for host side compilation we want to use the `__assert_fail()` from the glibc library. Tested the code by compiling with below steps ``` python3 tools/amd_build/build_amd.py python3 setup.py develop --cmake-only cmake -DHIP_HIPCC_FLAGS_RELEASE="-DNDEBUG" build cmake --build build ``` The UT test_fixed_cuda_assert_async is still skipped due performance overhead. cc @jithunnair-amd Pull Request resolved: https://github.com/pytorch/pytorch/pull/81790 Approved by: https://github.com/shintaro-iwasaki, https://github.com/jeffdaily, https://github.com/malfet	2022-08-16 19:22:31 +00:00
Zachary DeVito	4128712397	Propagate CUDAOutOfMemoryError to Python. (#83146 ) The intention is to make it easier to catch this situation for debugging, logging, or application-specific recovery. Pull Request resolved: https://github.com/pytorch/pytorch/pull/83146 Approved by: https://github.com/albanD	2022-08-11 21:32:11 +00:00
Zachary DeVito	726d040692	annotated allocator snapshots (#82146 ) Record stack trace information for each allocated segment in the allocator. It takes around 1.5us to record 50 stack frames of context. Since invoking a Pytorch operator is around 8us, this adds minimal overhead but we still leave it disabled by default so that we can test it more on real workloads first. Stack information is kept both for allocated blocks and the last allocation used inactive blocks. We could potential keep around the _first_ allocation that caused the block to get allocated from cuda as well. Potential Followups: * stack frame entries are small (16 bytes), but the list of Frames is not compressed eventhough most frames will share some entries. So far this doesn't produce huge dumps (7MB for one real workload that uses all memory on the GPU), but it can be much smaller through compression. * Code to format the information is slow (a few seconds) because it uses python and FlameGraph.pl * Things allocated during the backward pass have no stack frames because they are run on another C++ thread. Pull Request resolved: https://github.com/pytorch/pytorch/pull/82146 Approved by: https://github.com/albanD	2022-08-09 17:21:35 +00:00
Aidyn-A	da0a3fe058	[Re-land] [CUDA graphs] Clear autocast amp cache (#81896 ) Re-lands #81558 that got reverted due failing tests. This failure happened because of the test that I poorly designed. [The loop here](https://github.com/pytorch/pytorch/pull/81558/files#diff-893b1eea27352f336f4cd832919e48d721e4e90186e63400b8596db6b82e7450R3837) is doing `cache_enabled=False` and then `cache_enabled=True`. By doing this loop the graph from previous iteration (case `False`) conflicts with the next one (case `True`). I redesigned the test such that it does not do any loops. The new test does separate function calls with different argument values. Pull Request resolved: https://github.com/pytorch/pytorch/pull/81896 Approved by: https://github.com/ngimel	2022-08-02 23:22:00 +00:00
Kurt Mohler	14d0296e5c	Rename `_Typed/_UntypedStorage` to `Typed/UntypedStorage` and update docs (#82438 ) ### Description Since the major changes for `_TypedStorage` and `_UntypedStorage` are now complete, they can be renamed to be public. `TypedStorage._untyped()` is renamed to `TypedStorage.untyped()`. Documentation for storages is improved as well. ### Issue Fixes #82436 ### Testing N/A Pull Request resolved: https://github.com/pytorch/pytorch/pull/82438 Approved by: https://github.com/ezyang	2022-07-30 19:37:08 +00:00
Eddie Yan	0b2566456f	[CUDNN] Update tests and dispatching for CUDNN V8 API behavior for `bfloat16` convs (#81139 ) cuDNN via the V8 API supports `bfloat16` on Ampere (`>= (8, 0)` but not older devices) which might be unexpected given current test settings. This PR fixes some dispatching to check the device capability before dispatching `bfloat16` convs and adjusts the expected failure conditions for the autocast test. CC @xwang233 @ptrblck Pull Request resolved: https://github.com/pytorch/pytorch/pull/81139 Approved by: https://github.com/ngimel	2022-07-29 23:28:58 +00:00
PyTorch MergeBot	f5b460b200	Revert "[CUDA graphs] Clear autocast amp cache (#81558 )" This reverts commit `e9d07bd4f0`. Reverted https://github.com/pytorch/pytorch/pull/81558 on behalf of https://github.com/janeyx99 due to Breaks windows 11.6 tests on trunk `e9d07bd4f0`	2022-07-21 12:46:36 +00:00
Aidyn-A	e9d07bd4f0	[CUDA graphs] Clear autocast amp cache (#81558 ) According to [autocast_mode.cpp](https://github.com/pytorch/pytorch/blob/master/aten/src/ATen/autocast_mode.cpp) `cached_casts` is to be cleared at the end of each forward pass. However, this was not the case in current implementation of `make_graphed_callables` so a graph created the following way: ``` with torch.cuda.amp.autocast(cache_enabled=True): graphed_foo = torch.cuda.make_graphed_callables(foo, tensors) ``` Behaves incorrectly. cc @ptrblck Pull Request resolved: https://github.com/pytorch/pytorch/pull/81558 Approved by: https://github.com/ngimel	2022-07-21 01:44:14 +00:00
Jeff Daily	ff6655defb	[ROCm] unskip external streams tests (#80922 ) These two tests are passing for ROCm 5.1.1 and 5.2. Pull Request resolved: https://github.com/pytorch/pytorch/pull/80922 Approved by: https://github.com/cpuhrsch	2022-07-08 21:29:29 +00:00
Nikita Shulga	1ad7ef3f21	Add check for cuda lazy init (#80912 ) Validate that no CUDA calls are made during `import torch` call, by importing torch and limited visible devices to non-existing device Should prevent regressions like ones reported in https://github.com/pytorch/pytorch/issues/80876 Pull Request resolved: https://github.com/pytorch/pytorch/pull/80912 Approved by: https://github.com/ngimel, https://github.com/atalman	2022-07-06 01:39:27 +00:00
Jeff Daily	20d56d2b32	increase sleep for TestCuda.test_caching_pinned_memory_multi_gpu (#76601 ) Fixes #68299. Fixes #70875. Test is flaky on ROCm because the HIP runtime occasionally copies asynchronously too quickly for the current sleep value of 50ms. This is not a bug. Increasing the sleep value to 1s to avoid flakiness. Pull Request resolved: https://github.com/pytorch/pytorch/pull/76601 Approved by: https://github.com/pruthvistony, https://github.com/malfet	2022-06-14 21:10:35 +00:00
Michael Carilli	ba27ee9e8f	[CUDA graphs] Allows Adam and AdamW to be capture-safe (#77862 ) Near term fix for https://github.com/pytorch/pytorch/issues/76368. Q. Why does the user need to request `capturable=True` in the optimizer constructor? Why can't capture safety be completely automatic? A. We need to set up capture-safe (device-side) state variables before capture. If we don't, and step() internally detects capture is underway, it's too late: the best we could do is create a device state variable and copy the current CPU value into it, which is not something we want baked into the graph. Q. Ok, why not just do the capture-safe approach with device-side state variables all the time? A. It incurs several more kernel launches per parameter, which could really add up and regress cpu overhead for ungraphed step()s. If the optimizer won't be captured, we should allow step() to stick with its current cpu-side state handling. Q. But cuda RNG is a stateful thing that maintains its state on the cpu outside of capture and replay, and we capture it automatically. Why can't we do the same thing here? A. The graph object can handle RNG generator increments because its capture_begin, capture_end, and replay() methods can see and access generator object. But the graph object has no explicit knowledge of or access to optimizer steps in its capture scope. We could let the user tell the graph object what optimizers will be stepped in its scope, ie something like ```python graph.will_use_optimizer(opt) graph.capture_begin() ... ``` but that seems clunkier than an optimizer constructor arg. I'm open to other ideas, but right now I think constructor arg is necessary and the least bad approach. Long term, https://github.com/pytorch/pytorch/issues/71274 is a better fix. Pull Request resolved: https://github.com/pytorch/pytorch/pull/77862 Approved by: https://github.com/ezyang	2022-06-13 01:56:47 +00:00
Kurt Mohler	aea6e2c396	Merge torch.cuda._UntypedStorage into torch._UntypedStorage (#75459 ) Fixes #74933 Pull Request resolved: https://github.com/pytorch/pytorch/pull/75459 Approved by: https://github.com/ezyang	2022-05-19 13:54:39 +00:00
Michael Carilli	929f1d5317	[RELAND] Adds torch.cuda.is_current_stream_capturing (#77789 ) Resubmit of https://github.com/pytorch/pytorch/pull/77673, which was reverted due to Windows test failures: https://github.com/pytorch/pytorch/pull/77673#issuecomment-1130425845. I suspect these failures happened because I don't explicitly set a side stream for graph capture in the new test. Not setting a side stream explicitly is alright on Linux because cuda tests implicitly use a side stream. I think Windows cuda tests implicitly use the default stream, breaking capture and leaving the backend in a bad state. Other graphs tests explicitly set side streams and don't error in Windows builds, so i'm 95% sure doing the same for the new test will work. Pull Request resolved: https://github.com/pytorch/pytorch/pull/77789 Approved by: https://github.com/ezyang	2022-05-18 23:18:53 +00:00
Jeff Daily	de86146c61	rocblas alt impl during backward pass only (#71881 ) In preparation of adopting future rocblas library options, it is necessary to track when the backward pass of training is executing. The scope-based helper class `BackwardPassGuard` is provided to toggle state. Pull Request resolved: https://github.com/pytorch/pytorch/pull/71881 Approved by: https://github.com/albanD	2022-05-18 19:42:58 +00:00
PyTorch MergeBot	0d8a0f186b	Revert "Adds torch.cuda.is_current_stream_capturing (#77673 )" This reverts commit `d03d43df52`. Reverted https://github.com/pytorch/pytorch/pull/77673 on behalf of https://github.com/suo	2022-05-18 19:31:49 +00:00
Michael Carilli	d03d43df52	Adds torch.cuda.is_current_stream_capturing (#77673 ) Exposes a way to query if CUDA graph capture is underway on the current stream. Pull Request resolved: https://github.com/pytorch/pytorch/pull/77673 Approved by: https://github.com/ezyang	2022-05-18 16:46:35 +00:00
Eddie Yan	76b952bb35	[CUBLAS][TF32] Skip `test_cublas_allow_tf32_get_set` if `TORCH_ALLOW_TF32_CUBLAS_OVERRIDE` is set (#77298 ) Follow-up to #77114 to prevent test breakages when the environment variable is set. CC @xwang233 @ngimel @ptrblck Pull Request resolved: https://github.com/pytorch/pytorch/pull/77298 Approved by: https://github.com/xwang233, https://github.com/ngimel	2022-05-17 21:57:09 +00:00
Eddie Yan	e838137b3e	Add high level control of fp32 matmul precision; disable TF32 for matmuls by default #76440 CC @mruberry @ptrblck Pull Request resolved: https://github.com/pytorch/pytorch/pull/76509 Approved by: https://github.com/ngimel	2022-05-04 20:40:13 +00:00
Felipe Petroski Such	b0c5fba967	[CUDA Graphs] Fix OOM inside graph capture_begin release_cached_blocks calls this: ``` void synchronize_and_free_events() { TORCH_INTERNAL_ASSERT(captures_underway == 0); ``` Which means we can't call that function when we are capturing a cuda graph: ``` import torch with torch.cuda.graph(torch.cuda.CUDAGraph()): torch.zeros(2 ** 40, device="cuda") ``` results in: ``` RuntimeError: captures_underway == 0INTERNAL ASSERT FAILED at "/tmp/torch/c10/cuda/CUDACachingAllocator.cpp":1224, please report a bug to PyTorch. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/76247 Approved by: https://github.com/ngimel	2022-04-29 17:42:04 +00:00
Jeff Daily	e846ef8818	add rocm ciflow/slow workflow Enables additional tests that historically have been missed for ROCm CI. Pull Request resolved: https://github.com/pytorch/pytorch/pull/72686 Approved by: https://github.com/seemethere	2022-04-22 17:41:28 +00:00
Ivan Yashchuk	4bb5e6e830	Fix `test_reduce_add_coalesced` failure (#74027 ) Summary: Recent change (https://github.com/pytorch/pytorch/pull/69751) introduced the requirement of using `.coalesce()` explicitly in the tests. Unfortunately, not all tests are run in the current CI configuration and one test failure slipped through. Fixes https://github.com/pytorch/pytorch/issues/74015. Pull Request resolved: https://github.com/pytorch/pytorch/pull/74027 Reviewed By: samdow Differential Revision: D34858112 Pulled By: mruberry fbshipit-source-id: 8904fac5e2b5335684a21f95a22646469478eb81 (cherry picked from commit 06d6e6d2a796af0e8444f4c57841a07ec4f67c9f)	2022-03-15 06:29:54 +00:00
Michael Carilli	2f957f513e	Deletes unused line in test_autocast_rnn (#73195 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/73195 Reviewed By: mruberry Differential Revision: D34557677 Pulled By: ngimel fbshipit-source-id: 284018b4596471332d0e90a08e2c38303fb2b3ae (cherry picked from commit bbf6913009e206c02e124c49ab80ef9596f7fcad)	2022-03-02 01:27:55 +00:00
Shintaro Iwasaki	7dc2cfa249	[c10][rocm] fix __assert_fail() declaration mismatch error (#73040 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/73040 This patch fixes a compilation error in PyTorch with ROCm when `NDEBUG` is passed. ## Problem Forward declaration of `__host__ __device__ __assert_fail()` is used in `c10/macros/Macros.h` for HIP compilation when `NDEBUG` is set However, HIP has `__device__ __assert_fail()` in `hip/amd_detail/amd_device_functions.h`, causing a function type error. This issue does not appear in ROCm CI tests since it happens only when `NDEBUG` is passed. ## Solution [EDIT] After the discussion on GitHub, we chose to entirely disable `CUDA_KERNEL_ASSERT()` for ROCm. --- To solve this compilation error, this patch disables `CUDA_KERNEL_ASSERT()`, which uses `__assert_fail()` when 1. `c10/macros/Macros.h` is included for `.hip` (precisely speaking, `__HIP__` or `__HIP_ARCH__` is defined), and 2. `NDEBUG` is passed. Note that there's no impact on default compilation because, without a special compilation flag, those HIP files are compiled without `-NDEBUG`. And that's why this issue has not been found. ### Justification [1] We cannot declare one host-and-device function for two separate host and device functions. ``` __device__ int func() {return 0}; __host__ int func() {return 0}; // Compile error (hipcc) // __device__ __host__ int func(); ``` [2] Forward declaration of a correct `__device__` only `__assert_fail()` for `__HIP__` causes the following error: ``` pytorch/c10/util/TypeCast.h:135:7: error: reference to __device__ function '__assert_fail' in __host__ __device__ function ERROR_UNSUPPORTED_CAST ^ pytorch/c10/util/TypeCast.h:118:32: note: expanded from macro 'ERROR_UNSUPPORTED_CAST' #define ERROR_UNSUPPORTED_CAST CUDA_KERNEL_ASSERT(false); ^ pytorch/c10/macros/Macros.h:392:5: note: expanded from macro 'CUDA_KERNEL_ASSERT' __assert_fail( ``` [3] Maybe there's a way to properly define `__assert_fail()` for HIP + NDEBUG, but this might be too much. Please let me just disable it. ### Technical details Error ``` pytorch/c10/macros/Macros.h:368:5: error: __host__ __device__ function '__assert_fail' cannot overload __device__ function '__assert_fail' __assert_fail( ^ /opt/rocm/hip/include/hip/amd_detail/amd_device_functions.h:1173:6: note: previous declaration is here void __assert_fail(const char assertion, ``` CUDA definition (9.x) of `__assert_fail()` ``` #elif defined(__GNUC__) extern __host__ __device__ __cudart_builtin__ void __assert_fail( const char , const char , unsigned int, const char ) __THROW; ``` ROCm definition (the latest version) ``` // `2b59661f3e/include/hip/amd_detail/amd_device_functions.h (L1172-L1177)` extern "C" __device__ __attribute__((noinline)) __attribute__((weak)) void __assert_fail(const char assertion, const char file, unsigned int line, const char function); ``` Test Plan: CI + reproducer ``` python3 tools/amd_build/build_amd.py python3 setup.py develop --cmake-only cmake -DHIP_HIPCC_FLAGS_RELEASE="-DNDEBUG" build cmake --build build ``` Reviewed By: xw285cornell Differential Revision: D34310555 fbshipit-source-id: 7542288912590533ced3f20afd2e704b6551991b (cherry picked from commit 9e52196e36820abe36bf6427cabc7389d3ea6cb5)	2022-03-01 04:35:30 +00:00
Philip Meier	b5f2574f36	no longer coalesce sparse COO tensors before comparison (#69751 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/69751 cc nikitaved pearu cpuhrsch IvanYashchuk Test Plan: Imported from OSS Reviewed By: zou3519 Differential Revision: D34262453 Pulled By: ezyang fbshipit-source-id: e2e62d2aa03fc569d2951c880960b256f5dc4aaa (cherry picked from commit `cb6b0ef719`)	2022-02-17 02:33:08 +00:00
Kurt Mohler	8e7fe87630	Rename `Typed/UntypedStorage` to `_Typed/_UntypedStorage` (#72540 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/72540 Reviewed By: jbschlosser Differential Revision: D34216823 Pulled By: bdhirsh fbshipit-source-id: 1bc9930ab582771ebf02308e035576cd1a0dbe47 (cherry picked from commit `329238f612`)	2022-02-15 23:53:01 +00:00
Louis Feng	83b3b5fb00	[PyTorch] Support NVTX range_start and range_end (#70030 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/70030 range_push and range_pop do not support multi-thread. It only works for push and pop range in the same thread. For process level ranges, we should use range_start and range_end. This is important because PyTorch forward is on one thread, while the autograd is on a different thread. See NVidia implementation documentation: `cab2dec760/NSight/nvToolsExt.h (L397-L407)` Test Plan: ``` buck test caffe2/test:cuda Started reporting to test run: https://www.internalfb.com/intern/testinfra/testrun/8162774391483460 ✓ ListingSuccess: caffe2/test:cuda - main (19.640) Summary ListingSuccess: 1 If you need help understanding your runs, please follow the wiki: https://fburl.com/posting_in_tpx_users Finished test run: https://www.internalfb.com/intern/testinfra/testrun/8162774391483460 ``` Reviewed By: malfet Differential Revision: D33155244 fbshipit-source-id: c7d5143f6da9b6ef0e0811e2fcae03a3e76f24de (cherry picked from commit `22134e91b7`)	2022-02-07 17:31:57 +00:00
Andrew Tulloch	0099796978	[CUDA Pinned Memory] [Retry] Alternative implementation of pinned memory allocator focusing on multi-threaded scalability (#69299 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/69299 https://github.com/pytorch/pytorch/pull/68906 + https://github.com/pytorch/pytorch/pull/68749 plugged one correctness hole (non-blocking copies of offset pinned memory tensors) while introducing another (non-blocking copies of pinned memory tensors with a non-standard DataPtr context). In this revision, we use both the tensor data pointer and context to attempt to identify the originating block in the pinned memory allocator. Test Plan: New unit tests added to cover the missing case previously. Reviewed By: yinghai Differential Revision: D32787087 fbshipit-source-id: 0cb0d29d7c39a13f433eb1cd423dc0d2a303c955 (cherry picked from commit `297157b1a1`)	2022-01-27 01:33:55 +00:00
Mike Ruberry	e0d829a266	Kill the test_torch.py mixin and creates test_scatter_gather_ops (#71691 ) Summary: Per title. Also annotates test_torch.py with additional cleanup tasks and adds empty sample inputs to elementwise unary and binary OpInfos. Pull Request resolved: https://github.com/pytorch/pytorch/pull/71691 Reviewed By: ngimel Differential Revision: D33735126 Pulled By: mruberry fbshipit-source-id: 8cc097a7581a8b620540c95b2a5889c1165ecf23 (cherry picked from commit `5c6a245a3f`)	2022-01-24 09:32:32 +00:00
Leo Fang	67941c8a94	Document `torch.cuda.ExternalStream`, `torch.cuda.caching_allocator_alloc` and `torch.cuda.caching_allocator_delete` (#70126 ) Summary: Fixes https://github.com/pytorch/pytorch/issues/67414. Fixes https://github.com/pytorch/pytorch/issues/70117. cc brianjo mruberry ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/70126 Reviewed By: mruberry Differential Revision: D33542910 Pulled By: ngimel fbshipit-source-id: 4b870f4dceca6ee4cc8fba58819f1cb18ac9f857	2022-01-12 15:44:40 -08:00
Jane Xu	20489ebdc9	Increase tensor size for mem check tests (#70603 ) Summary: Fixes https://github.com/pytorch/pytorch/issues/70226 Pull Request resolved: https://github.com/pytorch/pytorch/pull/70603 Reviewed By: mruberry Differential Revision: D33410439 Pulled By: janeyx99 fbshipit-source-id: e94615ece6d0fdf230de5297118678b70f34a18c	2022-01-05 08:27:48 -08:00
Jane Xu	c555b7bacb	GHA: Remove caffe2 check in Windows shard 1 smoke tests (#70010 ) Summary: Windows shard 1 hasn't actually been running any tests because the script that does so exited before running the python tests but did not report an error. This has been happening to all windows tests across the board, for example https://github.com/pytorch/pytorch/runs/4526170542?check_suite_focus=true Removing the caffe2.python check passes the smoke tests now. You can observe that the run_test.py file is called in the windows cpu job now https://github.com/pytorch/pytorch/runs/4541331717?check_suite_focus=true Pull Request resolved: https://github.com/pytorch/pytorch/pull/70010 Reviewed By: malfet, seemethere Differential Revision: D33161291 Pulled By: janeyx99 fbshipit-source-id: 85024b0ebb3ac42297684467ee4d0898ecf394de	2021-12-20 16:05:38 -08:00
Mike Ruberry	84b7832010	Updates CUDA memory leak check to verify against driver API and print more diagnostic information (#69556 ) Summary: Per title Pull Request resolved: https://github.com/pytorch/pytorch/pull/69556 Reviewed By: mrshenli Differential Revision: D32954770 Pulled By: mruberry fbshipit-source-id: a6c2ae6f704422c178569980ca4b9c72c4272f55	2021-12-17 23:37:49 -08:00
Mike Ruberry	dc87cf5fe1	Fixes mem_get_info when querying on a device other than the current device (#69640 ) Summary: Also fixes the documentation failing to appear and adds a test to validate that op works with multiple devices properly. Pull Request resolved: https://github.com/pytorch/pytorch/pull/69640 Reviewed By: ngimel Differential Revision: D32965391 Pulled By: mruberry fbshipit-source-id: 4fe502809b353464da8edf62d92ca9863804f08e	2021-12-08 23:04:30 -08:00
Dennis van der Staay	cbe0a38d8c	Back out "[CUDA Pinned Memory] Event recording with non-blocking copies should track the storage context, not the tensor data pointer" (#69193 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/69193 Reviewed By: xing-liu, yuchenhao Differential Revision: D32748570 fbshipit-source-id: bd73d7567f94c70daeace49d4081381b8adf2d77	2021-12-01 19:30:08 -08:00
Andrew Tulloch	d44e610efa	[CUDA Pinned Memory] Event recording with non-blocking copies should track the storage context, not the tensor data pointer (#68749 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/68749 The logic for asynchronous copies (either HtoD or DtoH) using cudaMemcpyAsync relies on recording an event with the caching host allocator to notify it that a given allocation has been used on a stream - and thus it should wait for that stream to proceed before reusing the host memory. This tracking is based on the allocator maintaining a map from storage allocation pointers to some state. If we try to record an event for a pointer we don't understand, we will silently drop the event and ignore it (`9554ebe44e/aten/src/ATen/cuda/CachingHostAllocator.cpp (L171-L175)`). Thus, if we use the data_ptr of a Tensor instead of the storage allocation, then reasonable code can lead to incorrectness due to missed events. One way this can occur is simply by slicing a tensor into sub-tensors - which have different values of `data_ptr()` but share the same storage, for example: ``` image_batch = torch.randn(M, B, C, H, W).pin_memory() for m in range(M): sub_batch = image_batch[m].cuda(non_blocking=True) # sub_batch.data_ptr() != image_batch.data_ptr() except for m == 0. # however, sub_batch.storage().data_ptr() == image_batch.storage().data_ptr() always. ``` Therefore, we instead use the storage context pointer when recording events, as this is the same state that is tracked by the caching allocator itself. This is a correctness fix, although it's hard to determine how widespread this issue is. Using the storage context also allows us to use a more efficient structure internally to the caching allocator, which will be sent in future diffs. Test Plan: Test added which demonstrates the issue, although it's hard to demonstrate the race explicitly. Reviewed By: ngimel Differential Revision: D32588785 fbshipit-source-id: d87cc5e49ff8cbf59052c3c97da5b48dd1fe75cc	2021-11-24 13:20:22 -08:00
eqy	790763b0fe	Add an option to disable reduced precision reductions for FP16 GEMM (#67946 ) Summary: https://github.com/pytorch/pytorch/issues/67578 disabled reduced precision reductions for FP16 GEMMs. After benchmarking, we've found that this has substantial performance impacts for common GEMM shapes (e.g., those found in popular instantiations of multiheaded-attention) on architectures such as Volta. As these performance regressions may come as a surprise to current users, this PR adds a toggle to disable reduced precision reductions `torch.backends.cuda.matmul.allow_fp16_reduced_precision_reduction = ` rather than making it the default behavior. CC ngimel ptrblck stas00 Note that the behavior after the previous PR can be replicated with `torch.backends.cuda.matmul.allow_fp16_reduced_precision_reduction = False` Pull Request resolved: https://github.com/pytorch/pytorch/pull/67946 Reviewed By: zou3519 Differential Revision: D32289896 Pulled By: ngimel fbshipit-source-id: a1ea2918b77e27a7d9b391e030417802a0174abe	2021-11-09 17:27:20 -08:00
Jane Xu	2578de4851	[skip ci] Set test owner for test_cuda* tests (#66838 ) Summary: Action following https://github.com/pytorch/pytorch/issues/66232 cc ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/66838 Reviewed By: saketh-are Differential Revision: D31841411 Pulled By: janeyx99 fbshipit-source-id: 5cdffdef4a92f9adcef1143ae4598b052c5acc6b	2021-10-21 17:36:25 -07:00
arindamroy-eng	32e790997b	[Rocm]Reduce severity of detected possible memory leak from assertion to warning (#65973 ) Summary: Fixes https://github.com/pytorch/pytorch/issues/62533. In very rare cases, the decorator for detecting memory leak is throwing assertion, even when the test is passing, and the memory is being freed with a tiny delay. The issue is not being reproduced in internal testing, but shows up sometimes in CI environment. Reducing the severity of such detection to warning, so as not to fail the CI tests, as the actual test is not failing, rather only the check inside the decorator is failing. Limiting the change to ROCM only for now. cc jeffdaily sunway513 jithunnair-amd ROCmSupport Pull Request resolved: https://github.com/pytorch/pytorch/pull/65973 Reviewed By: anjali411 Differential Revision: D31776154 Pulled By: malfet fbshipit-source-id: 432199fca17669648463c4177c62adb553cacefd	2021-10-21 07:10:54 -07:00
Yanli Zhao	8173d4df69	move get_cycles_per_ms() to common_utils (#66798 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/66798 get_cycles_per_ms is copied and used in a few places, move it to common_utils so that it can be used as a shared util function ghstack-source-id: 140790599 Test Plan: unit tests Reviewed By: pritamdamania87 Differential Revision: D31706870 fbshipit-source-id: e8dccecb13862646a19aaadd7bad7c8f414fd4ab	2021-10-18 14:04:09 -07:00
Kurt Mohler	5883523c1d	Remove dtype from torch.Storage and use only torch.ByteStorage (#62030 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/62030 Remove dtype tracking from Python Storage interface, remove all the different `<type>Storage` classes except for `ByteStorage`, and update serialization accordingly, while maintaining as much FC/BC as possible Fixes https://github.com/pytorch/pytorch/issues/47442 * THE SERIALIZATION FORMAT IS FULLY FC/BC. We worked very hard to make sure this is the case. We will probably want to break FC at some point to make the serialization structure of tensors make more sense, but not today. * There is now only a single torch.ByteStorage class. Methods like `Tensor.set_` no longer check that the dtype of storage is appropriate. * As we no longer know what dtype of a storage is, we've removed the size method from Storage, replacing it with nbytes. This is to help catch otherwise silent errors where you confuse number of elements with number of bytes. * `Storage._new_shared` takes a `nbytes` kwarg and will reject previous positional only calls. `Storage._new_with_file` and `_set_from_file` require explicit element size arguments. * It's no longer possible to convert storages to different types using the float/double/etc methods. Instead, do the conversion using a tensor. * It's no longer possible to allocate a typed storage directly using FloatStorage/DoubleStorage/etc constructors. Instead, construct a tensor and extract its storage. The classes still exist but they are used purely for unpickling. * The preexisting serialization format stores dtype with storage, and in fact this dtype is used to determine the dtype of the tensor overall. To accommodate this case, we introduce a new TypedStorage concept that exists only during unpickling time which is used to temporarily store the dtype so we can construct a tensor. If you overrode the handling of pickling/unpickling, you MUST add handling for TypedStorage or your serialization code will degrade to standard file-based serialization. Original pull request: https://github.com/pytorch/pytorch/pull/59671 Reviewed By: soulitzer, ngimel Differential Revision: D29466819 Pulled By: ezyang fbshipit-source-id: 4a14e5d3c2b08e06e558683d97f7378a3180b00e	2021-10-05 13:50:34 -07:00
Michael Dagitses	b737629ff0	simplify op name determination into a single forward pass (#64261 ) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/64261 Note that this does not preserve byte-for-byte compatibility with existing names. Test Plan: * Rely on CI to catch gross errors. * Merge after release cut to catch subtle issues. Reviewed By: albanD Differential Revision: D30700647 Pulled By: dagitses fbshipit-source-id: 7b02f34b8fae3041240cc78fbc6bcae498c3acd4	2021-09-02 07:32:11 -07:00

1 2 3 4 5 ...

623 Commits