pytorch

mirror of https://github.com/zebrajr/pytorch.git synced 2025-12-06 12:20:52 +01:00

Author	SHA1	Message	Date
atalman	0d3d84d866	[CD] Windows Magma build 12.9 and cuda scripts (#155799 ) Scripts needed to build Magma and CUDA on windows Same as https://github.com/pytorch/pytorch/pull/146653 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155799 Approved by: https://github.com/jeanschmidt	2025-06-12 17:41:24 +00:00
Chris Sidebottom	430cc1c636	Run tests on Amazon EC2 M8g Instances (#153940 ) Requires machines configured here: https://github.com/pytorch/test-infra/pull/6642 This adds additional test runs against AWS Graviton4 processors, alongside existing runs against AWS Graviton3 and AWS Graviton2 processors. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153940 Approved by: https://github.com/fadara01, https://github.com/malfet	2025-06-12 17:33:08 +00:00
Shangdi Yu	522a18bd6c	Fix provenance unit test (#155747 ) Summary: Fix the test to adapt added provenance tracking in D75837494 Test Plan: ``` buck2 run @//mode/dev-nosan fbcode//caffe2/test:fx -- -r test_graph_provenance ``` Rollback Plan: Differential Revision: D76466778 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155747 Approved by: https://github.com/YUNQIUGUO	2025-06-12 17:26:43 +00:00
zpcore	50d8168c8b	[DTensor] Support in gradient placement for local_map() (#155181 ) Support `in_grad_placements` argument in torch.distributed.tensor.experimental.local_map(). The argument helps enforce placement of gradient of the input Dtensor. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155181 Approved by: https://github.com/wanchaol	2025-06-12 17:07:04 +00:00
henrylhtsang	6c0b42fd2f	[inductor][cutlass backend] Log prescreening elpase (#155508 ) Differential Revision: [D76311352](https://our.internmc.facebook.com/intern/diff/D76311352/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155508 Approved by: https://github.com/jingsh	2025-06-12 16:48:52 +00:00
Simon Mahns	c1ae768baa	Basic MTIA ATen CMake (#155477 ) Summary: Basic ATen CMake Differential Revision: D75203592 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155477 Approved by: https://github.com/andyanwang, https://github.com/cyyever	2025-06-12 16:29:32 +00:00
Laith Sakka	f4376cac54	unify symbolic_shapes and sizevars dynamic shapes APIs naming 1 (#154774 ) Inductor have a set of APIs that allows performing symbolic evaluations similar to that of symbolic shapes but it operates on sympy expressions instead of symnodes. Namings are not consistent making them consistent in this stack. Step 1 : unify statically_know_true naming! for consistent experience. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154774 Approved by: https://github.com/drisspg, https://github.com/bobrenjc93, https://github.com/eellison	2025-06-12 16:11:55 +00:00
Runtian (Rachel) Li	9df2e8020f	fix code indentation for fx.md (#155764 ) Fixes https://github.com/pytorch/pytorch/issues/155023 Related PR: #155482 Description: As discussed here https://github.com/pytorch/pytorch/pull/155482#pullrequestreview-2918032289, I removed indentation for python code blocks as a follow-up modification for fx.md Checklist: - [x] The issue being fixed is referenced above (Fixes https://github.com/pytorch/pytorch/issues/155023) - [x] Only one issue is addressed in this pull request - [x] Labels from the issue that this PR is fixing are added to this pull request - [x] No unnecessary issues are included into this pull request. @pytorchbot label "topic: docs" @pytorchbot label "topic: not user facing" @pytorchbot label docathon-h1-2025 @pytorchbot label module: docs Pull Request resolved: https://github.com/pytorch/pytorch/pull/155764 Approved by: https://github.com/svekars	2025-06-12 16:02:33 +00:00
Pian Pawakapan	75824035d3	[dynamic shapes] skip fused linear path if not definitely contiguous (#155051 ) Falls back to non-fused linear -> add bias path for non-contiguous tensors with unbacked sizes Pull Request resolved: https://github.com/pytorch/pytorch/pull/155051 Approved by: https://github.com/laithsakka	2025-06-12 15:55:21 +00:00
Catherine Lee	51560797ce	[CI] Reuse old whl: switch default to always (#155572 ) Switch default to always reuse old whl I have a few worries about API rate limits Pull Request resolved: https://github.com/pytorch/pytorch/pull/155572 Approved by: https://github.com/huydhn, https://github.com/malfet, https://github.com/seemethere, https://github.com/atalman	2025-06-12 15:43:29 +00:00
Aleksandar Samardžić	62fa3f5aeb	Support tuning of _grouped_mm (#153953 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153953 Approved by: https://github.com/ngimel	2025-06-12 15:39:35 +00:00
Henry Tsang	6b3eef6d31	[cutlass backend] Only consider to use re worker if nvcc doesn't exist (#155745 ) Differential Revision: D76463340 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155745 Approved by: https://github.com/masnesral	2025-06-12 15:23:52 +00:00
Manuel Candales	851a6fa82d	[MPS] Migrate softshrink (forward and backward) to Metal kernel (#155586 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155586 Approved by: https://github.com/malfet ghstack dependencies: #155304, #155316, #155462, #155479, #155571	2025-06-12 15:02:43 +00:00
PyTorch MergeBot	2a3b41cbd0	Revert "[CI] Use `setup-python` from for Mac tests (#155698 )" This reverts commit `2b9d638e33`. Reverted https://github.com/pytorch/pytorch/pull/155698 on behalf of https://github.com/malfet due to It causes weird flaky failures in MPS and do not upload usage logs anymore ([comment](https://github.com/pytorch/pytorch/pull/155698#issuecomment-2967120676))	2025-06-12 14:42:32 +00:00
Zhengxu Chen	0fd711df19	[export] Allow user frame to be None when symbolic shape tries to get stacktrace. (#155744 ) Summary: Fixing https://github.com/pytorch/pytorch/issues/155605 Test Plan: CI Rollback Plan: Differential Revision: D76463358 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155744 Approved by: https://github.com/angelayi	2025-06-12 14:36:29 +00:00
Bin Bao	dd1b6621bc	Remove C10_DEPRECATED references in c10 (#151058 ) Summary: Revive https://github.com/pytorch/pytorch/pull/138406. Only limit the scope to files in c10. Summary from the original PR, ``` Looking in the code I see // NB: __cplusplus doesn't work for MSVC, so for now MSVC always uses // the "__declspec(deprecated)" implementation and not the C++14 // "[[deprecated]]" attribute. We tried enabling "[[deprecated]]" for C++14 on // MSVC, but ran into issues with some older MSVC versions. But looking at the MSVC C++ support table I see that the [[deprecated]] attribute is supported as of MSVC 2015 and that the vast majority of C++17 features became supported in MSVC 2015 or later. Since PyTorch is C++17 now, I infer that PyTorch must not support versions of MSVC earlier than MSVC 2015, so the versions of MSVC supported by PyTorch must support [[deprecated]]. Therefore, since we are finished deprecating old MSVCs we can deprecate C10_DEPRECATED. ``` Test Plan: CI Differential Revision: D72762767 Pull Request resolved: https://github.com/pytorch/pytorch/pull/151058 Approved by: https://github.com/r-barnes	2025-06-12 13:38:03 +00:00
FFFrog	d632cf2cc9	[Easy][Code Clean] Remove the unused and undefined function in pickler (#155772 ) As the title stated. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155772 Approved by: https://github.com/malfet	2025-06-12 13:03:36 +00:00
Wang, Chuanqi	8e8d4b13b0	[XPU] Simplify XPU `make triton` by install from PyTorch source (#155675 ) Remove install from source code build Pull Request resolved: https://github.com/pytorch/pytorch/pull/155675 Approved by: https://github.com/atalman	2025-06-12 13:02:23 +00:00
David Berard	132babe7e0	[user triton] dynamo support for new host-side TMA API (#155662 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155662 Approved by: https://github.com/aakhundov ghstack dependencies: #155510	2025-06-12 12:56:23 +00:00
Aaron Gokaslan	9cced33c7c	[BE]: Update cudnn to 9.10.2.21 (#155576 ) Update to CUDNN 9.10.2.21 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155576 Approved by: https://github.com/eqy, https://github.com/atalman	2025-06-12 12:50:36 +00:00
atalman	c199a4d0fd	Move non inductor workflows cuda 12.6->cuda 12.8 (#155234 ) Move non inductor workflows cuda 12.6->cuda 12.8 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155234 Approved by: https://github.com/Skylion007, https://github.com/zxiiro, https://github.com/cyyever, https://github.com/malfet	2025-06-12 12:42:34 +00:00
leeeizhang	eecaa0bbc6	[Multiprocesing] Fix `_release_ipc_counter` missing in rebuilding cuda ipc tensor with UntypedStorage (#155312 ) Fixes https://github.com/pytorch/pytorch/issues/155311 To avoid `torch.multiprocessing.reductions::rebuild_cuda_tensor` failed on untyped storage, this FIX PR adds the `_release_ipc_counter` into UntypedStorage like the previous legacy typed storage. `e2d141dbde/torch/storage.py (L1466-L1469)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155312 Approved by: https://github.com/mikaylagawarecki	2025-06-12 10:41:58 +00:00
Laith Sakka	0029259bdf	Add view_simple as meta function for view, and avoid calling reshape_view_helper. (#154757 ) address https://github.com/pytorch/pytorch/issues/153303 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154757 Approved by: https://github.com/bobrenjc93, https://github.com/leslie-fang-intel	2025-06-12 09:58:15 +00:00
Michael Lazos	d3d655ad14	[Hierarchical-Compile] Hash int args in addition to input shapes (#155655 ) Fixes Swsl_resnext101_32x16d in TIMM Pull Request resolved: https://github.com/pytorch/pytorch/pull/155655 Approved by: https://github.com/anijain2305	2025-06-12 06:35:12 +00:00
David Berard	c3ecabf059	[inductor][triton pin] add support for new TMA API for mm.py templates (#155723 ) Triton 3.4 will remove the experimental TMA APIs: https://github.com/triton-lang/triton/pull/6488 For mm.py templates, this PR adds support for using the new APIs when they are available (and otherwise falls back to the experimental APIs). For flex_attention, we'll remove TMA support for Triton 3.2 and 3.3 (versions of triton that don't have the new API). For mm_scaled_grouped.py, https://github.com/pytorch/pytorch/pull/150944 will remove TMA support for Triton 3.2. Note: we attempted this earlier with https://github.com/pytorch/pytorch/pull/154858, but this broke TMA usage in Triton 3.2. Differential Revision: [D76444471](https://our.internmc.facebook.com/intern/diff/D76444471) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155723 Approved by: https://github.com/NikhilAPatel	2025-06-12 06:25:47 +00:00
Nikita Shulga	2b9d638e33	[CI] Use `setup-python` from for Mac tests (#155698 ) Instead of `setup-miniconda` - Remove `CONDA_RUN` macro... - Hack the search path in `macos-test.sh` to put both python and python3 aliases first in the path (not sure what other action are messing with path environment variable) - Skip `TestMultiprocessing.test_fs_sharing` as even though it completes, it hangs on the shutdown both in CI and in all local setups I have - Skip `TestCppExtensionOpenRgistration.test_base_device_registration` as it hangs on the shutdown as well Pull Request resolved: https://github.com/pytorch/pytorch/pull/155698 Approved by: https://github.com/atalman ghstack dependencies: #155476, #155493, #155601, #155515, #155697	2025-06-12 04:58:00 +00:00
Yiming Zhou	57e4d7b5cc	[nativert] Move DelegateExecutor to PyTorch core (#155581 ) Summary: Moves DelegateExecutor base class to PyTorch core. It provides the extension point of backend delegation for NativeRT. Torch Native Runtime RFC: pytorch/rfcs#72 Test Plan: This is only a virtual base class. So relying on internal CI is sufficient. Rollback Plan: Differential Revision: D76351984 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155581 Approved by: https://github.com/zhxchen17	2025-06-12 04:33:31 +00:00
Animesh Jain	a9d5157e25	[dynamo] Use BINARY_SUBSCR for pre-graph bytecode for regular dict accesses (#155727 ) vLLM profiler sets with_stack=True that shows the dict_getitem on the profiler, both inflating the numbers and confusing compile users. This PR keeps BINARY_SUBSCR for regular dicts, while using `dict.__getitem__` only for dict subclasses. Using binary_subscr is little bit faster, but not enough to make any major latency improvements. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155727 Approved by: https://github.com/zou3519, https://github.com/StrongerXi, https://github.com/jansel	2025-06-12 04:02:29 +00:00
Animesh Jain	c9e9a0c823	[inductor][invoke_subgraph] Mark invoke_subgraph outputs as user_visible to constrain output strides (#155395 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155395 Approved by: https://github.com/zou3519	2025-06-12 03:58:16 +00:00
Shangdi Yu	9f5153b1a4	Preserve GrpahModule node stack trace after torch package deserializaion re-tracing (#155638 ) Summary: urrently the node.meta["stack_trace"] is not preserved when we torch package/load GraphModule, which means the original stack trace is lost. When we re-trace the packaged graph module, we just get a stack trace like fx-generated._0...... Adding the node.meta["stack_trace"] to torch packaged graph module Test Plan: ``` buck2 run @//mode/dev-nosan fbcode//caffe2/test:package -- -r TestPackageFX ``` Rollback Plan: Differential Revision: D76379692 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155638 Approved by: https://github.com/angelayi	2025-06-12 03:48:27 +00:00
Nikita Shulga	ce9ba071fd	[BE] Fix warning in open_registration_extension.cpp (#155755 ) Namely ``` /Users/nshulga/git/pytorch/pytorch/test/cpp_extensions/open_registration_extension.cpp:306:33: warning: left operand of comma operator has no effect [-Wunused-value] 306 \| at::Tensor first = at::empty((2,3)).to(at::DeviceType::PrivateUse1); ``` Or switching between Python and C++ is hard In Python `(2, 3)` creates a tuple, in C/C++ it's just a integral literal 3 P.S. I could have vibe-coded the fix with Claude: https://claude.ai/share/82479e88-84cb-4299-aa2f-dafd28ee2d55 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155755 Approved by: https://github.com/huydhn, https://github.com/atalman	2025-06-12 03:01:30 +00:00
Yiming Zhou	d96dec8415	[export] Fix serialization for call_torchbind hop with as_none argument (#155647 ) Summary: As title. D75251816 broke one internal test. This diff fixes it. Test Plan: Internal CI Differential Revision: D76383202 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155647 Approved by: https://github.com/ydwu4	2025-06-12 02:59:03 +00:00
Kazuaki Ishizaki	b00b641ff1	[Docs] Convert to markdown: accelerator.rst, amp.rst, autograd.rst, backends.rst, benchmark_utils.rst (#155762 ) Fixes #155013 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155762 Approved by: https://github.com/svekars	2025-06-12 02:55:06 +00:00
Xia, Weiwen	b6f84b3b0f	[Inductor][CPU] Use AMX-based microkernels when M > 4 for GEMM template for INT4 weight (#155444 ) Summary GEMM templates for INT4 weights are used for lowering `aten._weight_int4pack_mm_for_cpu` with Inductor when max-autotune is on. Currently, AMX-based microkernels are used only when M >= 16 if input tensor has shape [M, K]. However, we find that AMX kernel brings performance benefit when 4 < M < 16. For example, on a 6th gen of Intel(R) Xeon(R) platform, E2E latency can be improved by up to > 20% when running Llama-3.1-8B on 32 cores for M = 8. So, this PR changes the threshold so that AMX is used when M > 4. Test plan ``` pytest test/inductor/test_cpu_select_algorithm.py -k test_int4_woq_mm ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155444 Approved by: https://github.com/sanchitintel, https://github.com/leslie-fang-intel	2025-06-12 02:28:48 +00:00
Simon Fan	212575f994	[ca] Annotate AccumulateGrad branching and add polyfill tests (#155289 ) Annotates AccumulateGrad and tracks the semantics for AccumulateGrad's polyfill , except for Scenario 1.4: Cloning MKLDNN new_grad and Scenario 2.2: Vmap-incompatible. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155289 Approved by: https://github.com/jansel, https://github.com/albanD	2025-06-12 02:10:52 +00:00
Yu, Guangye	d84efde3f0	Move _storage_Use_Count to be gerneric (#155451 ) # Motivation `torch._C._storage_Use_Count` should be a generic API that is not aware of device type. It is also used in `337cd7c53d/torchtune/training/_activation_offloading.py (L323)` to do some memory optimization. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155451 Approved by: https://github.com/albanD	2025-06-12 01:39:04 +00:00
PyTorch MergeBot	8372d0986a	Revert "[PT2][partitioners] Add aten.split to view_ops list (#155424 )" This reverts commit `e1db10e05a`. Reverted https://github.com/pytorch/pytorch/pull/155424 on behalf of https://github.com/clee2000 due to I think this broke inductor/test_cpu_repro.py::CPUReproTests::test_transpose_with_norm [GH job link](https://github.com/pytorch/pytorch/actions/runs/15596830833/job/43931044625) [HUD commit link](`e1db10e05a`) but idk how, reverting to see if it fixes the problem ([comment](https://github.com/pytorch/pytorch/pull/155424#issuecomment-2964717706))	2025-06-12 01:38:34 +00:00
Catherine Lee	9b122aab5d	Fix set per proc memory fraction when running tests (#155631 ) env setting needs to happen before pool creation for it to take effect In theory this should fix some OOMs and also cause some OOMs, but this PR is green so idk alt options: use initializer? Pull Request resolved: https://github.com/pytorch/pytorch/pull/155631 Approved by: https://github.com/huydhn, https://github.com/malfet, https://github.com/seemethere, https://github.com/atalman	2025-06-12 01:28:08 +00:00
Pian Pawakapan	8ad6197b46	[draft export] avoid storing intermediate real tensors in proxies (#154630 ) Handles GC for non-strict draft export; GPU memory usage shouldn't be much more than eager mode + input tensors now. While trying to do draft export CPU offloading, I found out GC is feasible, because in non-strict, there's 2 places holding references to a `.real_tensor` attribute: 1) the FakeTensors in fake tensor prop, but these are held by the actual variables in the model's forward call, and so the real tensor gets gc-ed along with the fake one when the variable goes out of scope. 2) A clone of the fake tensor in 1) stored in `proxy.node.meta["val"]`, which was added in https://github.com/pytorch/pytorch/pull/150948. But we didn't actually need to store them on intermediate values; the placeholders are enough for retracing/lowering. Avoiding storing the intermediate values in 2), the values in 1) should be naturally GC-ed, and the real-tensor memory usage for non-strict should be pretty similar to eager computation? Strict still OOMs; dynamo still holds these in variable tracking, and not sure how to GC those. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154630 Approved by: https://github.com/angelayi, https://github.com/yushangdi	2025-06-12 01:18:57 +00:00
Shangdi Yu	4e19477196	[nativert] Move Pytree (#155136 ) Summary: fbcode/sigmoid/core/common -> fbcode/caffe2/torch/nativert/common Torch Native Runtime RFC: https://github.com/pytorch/rfcs/pull/72 Test Plan: ``` buck run fbcode//mode/dev-nosan //caffe2/test/cpp/nativert:pytree_test ``` OSS CI Rollback Plan: Differential Revision: D75965059 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155136 Approved by: https://github.com/zhxchen17, https://github.com/XuehaiPan, https://github.com/zou3519	2025-06-12 01:10:34 +00:00
Wanchao Liang	ee5c2908cb	[dtensor] refactor PlacementStrategy -> OpSpec, move utils to OpSchema (#155592 ) as titled. It's sometimes confusing to use PlacementStrategy as a name, as we also have OpStrategy and TupleStrategy, the latter two contain the former, so it is better to make the naming clearer. Renaming PlacementStrategy -> OpSpec as it is an operator spec that contains output_spec + input_specs. Also found some utils can be merged to OpSchema so included together in this PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/155592 Approved by: https://github.com/awgu	2025-06-12 00:51:36 +00:00
Huy Do	7485ef078f	Run torch.compile benchmark more frequently on H100 (#155719 ) We have more capacity now with 20+ `linux.aws.h100` runners, half of them are idle. Running benchmark more frequently would utilize these runner better and provide early signals multiple times per day. Running every 8 hours to start with. The workflow usually finishes within 5 hours https://github.com/pytorch/pytorch/actions/runs/15578331612/job/43878878434 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155719 Approved by: https://github.com/atalman	2025-06-12 00:24:21 +00:00
Ke Wen	9e9484d022	[SymmMem] Enable NVSHMEM for Triton (#155506 ) (This is an Experimental feature) Allow Triton kernels to invoke NVSHMEM device functions. ### Example Triton program Key parts: - Call `nvshmem.enable_triton()` to initialize; - Call `nvshmem.putmem_block` in Triton kernel; - Add `extern_libs` kwarg at kernel invocation. ``` import torch.distributed._symmetric_memory._nvshmem_triton as nvshmem @triton.jit def put_kernel( dst_ptr, src_ptr, numel: tl.constexpr, peer: tl.constexpr, BLOCK_SIZE: tl.constexpr, ): nvshmem.putmem_block(dst_ptr, src_ptr, numel, peer) if __name__ == "__main__": # Enable NVSHMEM for Triton nvshmem_lib = nvshmem.enable_triton() # Use torch Symmetric Memory to allocate Symmetric tensors ... peer = 1 - rank if rank == 0: kernel = put_kernel[(1, 1, 1)]( dst_ptr, src_ptr, numel=numel, peer=peer, BLOCK_SIZE=BLOCK_SIZE, extern_libs=nvshmem_lib, ) dist.barrier() if rank == 1: print(f"Rank {rank}: received {out=}") ``` ### Test output: ``` $ TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py -k test_triton_put Rank 0: writing value 5 to Peer 1 Rank 1: received out=tensor([5, 5, 5, 5, 5, 5, 5, 5], device='cuda:1', dtype=torch.int8) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155506 Approved by: https://github.com/ngimel, https://github.com/fegin, https://github.com/fduwjj	2025-06-12 00:22:49 +00:00
Justin Silver	cf9878d7a2	Fix #155022 rst to markdown conversion (#155540 ) Used [rst2myst tool](https://rst-to-myst.readthedocs.io/en/latest/) Fixes #155022 Docs comparison (check out the 'new' whenever docs build) 1. func.ux_limitations ([old](https://docs.pytorch.org/docs/main/func.ux_limitations.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155540/func.ux_limitations.html)) 2. func.whirlwind_tour ([old](https://docs.pytorch.org/docs/main/func.whirlwind_tour.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155540/func.whirlwind_tour.html)) 3. future_mod ([old](https://docs.pytorch.org/docs/main/future_mod.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155540/future_mod.html)) 4. futures ([old](https://docs.pytorch.org/docs/main/futures.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155540/futures.html)) 5. fx.experimental ([old](https://docs.pytorch.org/docs/main/fx.experimental.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155540/fx.experimental.html)) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155540 Approved by: https://github.com/AlannaBurke, https://github.com/svekars	2025-06-12 00:21:22 +00:00
Sidharth	7918978653	[dynamo] uploaded full json file of all unimplemented_v2() calls currently in repository (#155758 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155758 Approved by: https://github.com/williamwen42	2025-06-12 00:17:28 +00:00
Tsung-Hsien Lee	a6210fd07b	[c10d] Enhance `get_process_group_ranks()` to accept `group=None` (#154902 ) Summary: This diff enhances the `get_process_group_ranks()` function to accept `group=None` as an optional argument. This allows the function to return all ranks associated with the default process group if no group is specified. Test Plan: contbuild & OSS CI Rollback Plan: Differential Revision: D75817800 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154902 Approved by: https://github.com/wz337	2025-06-11 23:41:03 +00:00
eqy	bd3c32916c	[cuDNN] Enabled dilation for deterministic convolutions in cuDNN (#154292 ) Provides order-of-magnitude speedup over fallback impl. https://github.com/pytorch/pytorch/issues/28777 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154292 Approved by: https://github.com/Skylion007 Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>	2025-06-11 23:35:52 +00:00
Ankita George	c13e725edd	Updates to HFStorageReader to use TensorStorageMetadata instead of BytesStorageMetadata (#154518 ) As we prepare to support re-sharding, the current approach of using BytesStorageMetadata to read safetenstors won't work anymore. Before, we didn't need to read the metadata of the safetensors file from its header because we were just loading the contents of the file directly into tensors with safetensor.load() that would handle the metadata and deserialization. But now, in preparation of handling re-sharding, we need to read the metadata directly from the header of the safetensors file and store it directly in TensorStorageMetadata objects so that we can perform re-sharding. Re-sharding won't currently work, as we need extra metadata to be stored on each save, so that will be added in a subsequent PR. In addition this PR adds an integration test in addition to the unit tests. It also removes the HfFileSystem import because that's only needed if users are using HfFileSystem, but we want to support any backend. Differential Revision: [D74891998](https://our.internmc.facebook.com/intern/diff/D74891998/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154518 Approved by: https://github.com/saumishr	2025-06-11 23:35:05 +00:00
jafraustro	1b032384b1	Convert rst files to md (#155369 ) Fixes #155021 Fixes #155158 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155369 Approved by: https://github.com/svekars, https://github.com/malfet	2025-06-11 23:00:52 +00:00
Nikita Shulga	48921721d8	[MPS] Fix binary builds (#155733 ) Introduced by https://github.com/pytorch/pytorch/pull/155611 All functions in those headers must be static and inline Pull Request resolved: https://github.com/pytorch/pytorch/pull/155733 Approved by: https://github.com/seemethere, https://github.com/atalman	2025-06-11 22:55:33 +00:00

1 2 3 4 5 ...

88956 Commits