pytorch

mirror of https://github.com/zebrajr/pytorch.git synced 2025-12-06 12:20:52 +01:00

Author	SHA1	Message	Date
Jordan Zhu	4bdecd94ea	[modefile free][long tail] selectify fbcode/caffe2/defs.bzl (#148925 ) Summary: replace read_config with select For more info, please refer to the [doc](https://docs.google.com/document/d/1e0Hvht8WEHhcRvlCAodq_R9xnAtKBrAhdyvxcAqQjCw/edit?tab=t.hl8j18gza0cv) Test Plan: CI Reviewed By: malfet Differential Revision: D70267850 Pull Request resolved: https://github.com/pytorch/pytorch/pull/148925 Approved by: https://github.com/malfet	2025-04-28 16:04:28 +00:00
PyTorch MergeBot	9c864f9b0f	Revert "[Inductor UT] Generalize device-bias code in `test_flex_attention.py` (#151937 )" This reverts commit `4438400802`. Reverted https://github.com/pytorch/pytorch/pull/151937 on behalf of https://github.com/malfet due to Broke ASAN tests, probably by enabling too many tests https://hud.pytorch.org/hud/pytorch/pytorch/main/1?per_page=50&name_filter=asan&mergeEphemeralLF=true ([comment](https://github.com/pytorch/pytorch/pull/151937#issuecomment-2835151532))	2025-04-28 12:56:49 +00:00
PyTorch UpdateBot	0b6ea0b959	[xla hash update] update the pinned xla hash (#151210 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned xla hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151210 Approved by: https://github.com/pytorchbot	2025-04-28 11:45:09 +00:00
Anthony Shoumikhin	7cae7902a2	Add scripts to check xrefs and urls (#151844 ) Traverses the docs and code to find any broken links Pull Request resolved: https://github.com/pytorch/pytorch/pull/151844 Approved by: https://github.com/huydhn	2025-04-28 09:30:07 +00:00
Scott Wolchok	7e8b9b3f51	ReducedPrecisionFloatGemvFastPathKernel: Correctly type parallel_for lambda arguments as int64_t (#152233 ) This plus the previous irangeification PR seem like a better fix for #150637 than #150949 to me -- should make sure we are using 64-bit math for indexing everywhere. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152233 Approved by: https://github.com/Skylion007, https://github.com/cyyever ghstack dependencies: #152232	2025-04-28 07:19:26 +00:00
Scott Wolchok	3b7d6bbe8b	irangeify ReducedPrecisionFloatGemvKernel.cpp (#152232 ) We should be using irange, especially because we had 32-bit overflow issues in this file recently. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152232 Approved by: https://github.com/Skylion007	2025-04-28 07:19:26 +00:00
Gabriel Ferns	ce00ec7ecf	Enable max autotune for AOTInductor benchmark (#149309 ) With this PR, AOTinductor can choose to run into max-autotune mode when benchmarking. Pull Request resolved: https://github.com/pytorch/pytorch/pull/149309 Approved by: https://github.com/desertfire Co-authored-by: Gabriel Ferns <gabeferns@meta.com>	2025-04-28 06:54:26 +00:00
Nikita Shulga	13966d0bf5	[BE] Migrate dtype_abbrs into one location (#152229 ) Namely `torch.utils._dtype_abbrs.dtype_abbrs` Before that it was defined in various forms of completeness in `c02edba863/torch/fx/graph.py (L215)`, `c02edba863/torch/testing/_internal/common_utils.py (L5226)` and `c02edba863/torch/testing/_internal/logging_tensor.py (L17)` TODO: - Add linter that `torch.testing._internal` module is not referenced from any of the public facing APIs, as it can have extra dependencies such as `expect_test` Fixes https://github.com/pytorch/pytorch/issues/152225 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152229 Approved by: https://github.com/clee2000, https://github.com/Skylion007	2025-04-28 03:52:47 +00:00
Isalia20	899eec665c	[MPS] col2im kernel implementation (#152282 ) Fixes #151820 Also requested in #141287 Mainly based on the cuda kernel implementations Pull Request resolved: https://github.com/pytorch/pytorch/pull/152282 Approved by: https://github.com/malfet	2025-04-28 03:48:41 +00:00
Aart J.C. Bik	2503843673	Add check for 2-dim mask to COO mask computation (#151940 ) Follow up on discussion on https://github.com/pytorch/pytorch/pull/151794 Related to all fixes for https://github.com/pytorch/pytorch/issues/151351 Pull Request resolved: https://github.com/pytorch/pytorch/pull/151940 Approved by: https://github.com/Skylion007	2025-04-28 03:40:46 +00:00
Anatoly Myachev	4438400802	[Inductor UT] Generalize device-bias code in `test_flex_attention.py` (#151937 ) @EikanWang @etaf @guangyey please take a look Pull Request resolved: https://github.com/pytorch/pytorch/pull/151937 Approved by: https://github.com/liangan1, https://github.com/drisspg	2025-04-28 03:07:23 +00:00
Laith Sakka	98bd2bd1ab	Do not generate long log messages for suppressed data dependent errors. (#151023 ) TORCH_LOGS="all" python test/test_dynamic_shapes.py -k test_guard_or_true before: <img width="1065" alt="Screenshot 2025-04-10 at 9 55 27 AM" src="https://github.com/user-attachments/assets/3ee20de0-2902-4eb1-8ab0-80f1b974fb78" /> after: <img width="1124" alt="Screenshot 2025-04-10 at 9 54 35 AM" src="https://github.com/user-attachments/assets/4e7e1f0c-856c-417f-8763-bfe183e2450d" /> Note: we actually do not expect to see a log at all, this is an orthogonal issue in recording where it logs each error seen even when recording is not enabled? I will follow up with PR for that. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151023 Approved by: https://github.com/bobrenjc93	2025-04-28 00:39:52 +00:00
cyy	70d7638b0d	Fix clang-tidy suppression in torch/csrc/jit (#152271 ) Remove some clang-tidy suppression in torch/csrc/jit by applying fixes or refactoring. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152271 Approved by: https://github.com/Skylion007, https://github.com/malfet Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>	2025-04-27 21:18:39 +00:00
PyTorch MergeBot	c02edba863	Revert "Update OpenBLAS commit (#151547 )" This reverts commit `c4b0854750`. Reverted https://github.com/pytorch/pytorch/pull/151547 on behalf of https://github.com/malfet due to It breaks all aarch64 tests ([comment](https://github.com/pytorch/pytorch/pull/151547#issuecomment-2833593427))	2025-04-27 18:58:35 +00:00
cyy	b34146a093	Fix initGdsBindings declaration (#152277 ) Move initGdsBindings into the correct namespace. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152277 Approved by: https://github.com/Skylion007	2025-04-27 17:04:56 +00:00
Zizeng Meng	861945100e	[Kineto] Enable OOM observer (#152160 ) Summary: # Context: When memory leak happens, it usually trigger the OOM in the later iterations. The snapshot of full iteration will be huge and hard to interpret. On CUDA side, they provide OOM observer which generates snapshot when OOM happens with latest 1,500,000 entries for debugging. In this diff, we want to implement the feature on MTIA side Test Plan: Run this test with last diff in the stack. ``` buck run @//mode/opt kineto/libkineto/fb/mtia/integration_tests:mtia_memory_auto_trace_test ``` As shown, the memory_snapshot is generated when oom happens Log: P1794792326 Snapshot: https://fburl.com/pytorch_memory_visualizer/lx73y6s3 {F1977402355} Differential Revision: D71993315 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152160 Approved by: https://github.com/sraikund16	2025-04-27 15:56:44 +00:00
Aditya Tewari	c4b0854750	Update OpenBLAS commit (#151547 ) Motivation: Update OpenBLAS and change build script to enable SBGEMM kernels . Update pytorch `jammy` builds for aarch64 to use `install_openblas.sh` instead of `conda_install` Link to full [TorchInductor Performance Dashboard AArch64](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Wed%2C%2016%20Apr%202025%2009%3A35%3A26%20GMT&stopTime=Thu%2C%2017%20Apr%202025%2009%3A35%3A26%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cpu%20(aarch64)&lBranch=adi/update_openblas&lCommit=90701ab81bf61fd864d31e0aa7e88d97a1a8676c&rBranch=main&rCommit=40ce4fb24a536d175348df876f61956d4945778e) 1. This shows a promising speedup across most of the HF models in benchmark, specifically giving a significant boost to SDPA layers. 2. Overall torch-bench pass-rate increased `[87%, 65/75 → 96%, 72/75]` <img width="676" alt="Screenshot 2025-04-17 at 10 32 10" src="https://github.com/user-attachments/assets/a92dce0c-ecee-4466-8175-065df664dd71" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/151547 Approved by: https://github.com/malfet	2025-04-27 15:55:42 +00:00
Nikita Shulga	bb680b5a87	[MPSInductor] Fix masked_fill decomp (#152268 ) By adding `mps` to the list of accelerators that can work with CPU scalars Fixes `GPUTests.test_masked_fill_promotion_mps` Pull Request resolved: https://github.com/pytorch/pytorch/pull/152268 Approved by: https://github.com/kulinseth, https://github.com/dcci, https://github.com/Skylion007 ghstack dependencies: #152266	2025-04-27 15:50:46 +00:00
Yuanhao Ji	cbcf677223	[Dynamo] Replace `unimplemented` with `unimplemented_v2` in `torch/_dynamo/variables/lists.py` (#151873 ) Part of #147913 Replace `unimplemented` with`unimplemented_v2` in `torch/_dynamo/variables/lists.py` Pull Request resolved: https://github.com/pytorch/pytorch/pull/151873 Approved by: https://github.com/williamwen42 Co-authored-by: William Wen <william.wen42@gmail.com>	2025-04-27 11:59:45 +00:00
Yuanhao Ji	0423a7b322	[Dynamo] Replace `unimplemented` with `unimplemented_v2` in `torch/_dynamo/variables/nn_module.py` (#151895 ) Part of #147913 Replace `unimplemented` with`unimplemented_v2` in `torch/_dynamo/variables/nn_module.py` Pull Request resolved: https://github.com/pytorch/pytorch/pull/151895 Approved by: https://github.com/williamwen42 Co-authored-by: William Wen <william.wen42@gmail.com>	2025-04-27 11:54:42 +00:00
Anthony Shoumikhin	e2f9759bd0	Fix broken URLs (#152237 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152237 Approved by: https://github.com/huydhn, https://github.com/malfet	2025-04-27 09:56:42 +00:00
Nikita Shulga	cbcc03c2ad	[MPSInductor][BE] Only include headers when needed (#152266 ) Store headers used by shader in `MetalKernel.headers` Add headers when function depending on it gets invoked Generate majority of a special ops from template Delete two unused functors: `entr` and `xlog1py` as they are decomposed by inductor anyway Pull Request resolved: https://github.com/pytorch/pytorch/pull/152266 Approved by: https://github.com/Skylion007, https://github.com/jansel, https://github.com/dcci, https://github.com/cyyever	2025-04-27 05:09:50 +00:00
Bin Bao	a0d440a26a	[AOTI][reland] Remove typedef for half and bfloat16 (#151109 ) Summary: Reland https://github.com/pytorch/pytorch/pull/150657 typedef is prone to name collision. Explicitly spell out the actual aten types, needed for the libtorch-free codegen. Differential Revision: [D72878456](https://our.internmc.facebook.com/intern/diff/D72878456) Pull Request resolved: https://github.com/pytorch/pytorch/pull/151109 Approved by: https://github.com/angelayi	2025-04-26 23:17:35 +00:00
Zhiyi Zhang	225742838b	Add an additional check to trigger graph break for sparse tensor (#151897 ) Fixes #151522 This PR fixes the issue that Dynamo fails to trigger a graph break for sparse tensors in certain code paths. I added an additional check to handle this case, and it resolves the original problem. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151897 Approved by: https://github.com/jansel	2025-04-26 21:02:32 +00:00
Oguz Ulgen	e4a1a16bef	Check integrity of bytes in AppendingByteSerializer (#152139 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152139 Approved by: https://github.com/zou3519	2025-04-26 18:10:58 +00:00
co63oc	9480ed4cd3	Fix typos in multiple files (#152254 ) Fix typos in multiple files Pull Request resolved: https://github.com/pytorch/pytorch/pull/152254 Approved by: https://github.com/Skylion007	2025-04-26 17:18:39 +00:00
Aaron Gokaslan	6a62356857	[BE][Easy]: Change typing to DimsType in dim_reduction (#151677 ) Use prims_common DimsType to reduce duplication of DType Pull Request resolved: https://github.com/pytorch/pytorch/pull/151677 Approved by: https://github.com/albanD	2025-04-26 16:59:32 +00:00
Zhengxu Chen	203201255f	[dynamo] remove dead code for DATA_PTR_MATCH (#152206 ) Summary: Seems this guard is not created anywhere Test Plan: CI Differential Revision: D73682084 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152206 Approved by: https://github.com/anijain2305, https://github.com/jansel	2025-04-26 15:25:01 +00:00
Yukio Siraichi	ee8166e94f	Correctly handle duplicated arguments when merging input views. (#146275 ) Fix: #135099 This PR changes how we map the original inputs into the new set of inputs that take in the tensor input's base instead of their aliases. Problem: in order to create this mapping, we had a dictionary that mapped the hashed arguments into their respective indices. However, if there's a group of equal arguments, we will have only one mapping for such an argument. This breaks the assumption that there will be one mapping for each argument. Solution: map the hashed arguments into a list of indices. Then, we will be able to correctly reconstruct the parameters for the new calling convention. Pull Request resolved: https://github.com/pytorch/pytorch/pull/146275 Approved by: https://github.com/bdhirsh	2025-04-26 14:50:16 +00:00
FFFrog	580913290c	[Easy] The event_id of torch.cuda.Event and torch.xpu.Event always is 0 (#151226 ) Although torch.cuda.Event and torch.xpu.Event have cuda_event and sycl_event fields respectively, the event_id exposed from the base class torch.Event is always 0, which can confuse users. The memory of torch.Event is not useful to torch.cuda.Event and torch.xpu.Event, but we still need to inherit from torch.Event because CPython will check it. Repro with cuda: ``` >>> import torch >>> event = torch.cuda.Event() >>> event.cuda_event 0 >>> event.event_id 0 >>> event.record() >>> event.cuda_event 127982096 >>> event.event_id 0 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/151226 Approved by: https://github.com/albanD, https://github.com/guangyey ghstack dependencies: #151404, #151221, #151411	2025-04-26 14:18:22 +00:00
Davide Italiano	2ce9d2e9aa	[MPS/inductor] Adjust test_to_dtype_mps so that it works on the backend. (#152230 ) float64 isnt' supported for MPS, but we can still test the functionality with another type. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152230 Approved by: https://github.com/malfet, https://github.com/jansel	2025-04-26 13:54:53 +00:00
FFFrog	0f9b02c839	[Easy][torch.Event] Fix and improve the docs of torch.Event (#151411 ) Changes: - add detailed function or class signature - fix the wrong display of torch.Event.wait and torch.Event.record Pull Request resolved: https://github.com/pytorch/pytorch/pull/151411 Approved by: https://github.com/albanD ghstack dependencies: #151404, #151221	2025-04-26 13:52:38 +00:00
FFFrog	bd7dc1b17d	[Easy] Fix the function signature of torch.Event (#151221 ) As the title stated. The difference between declaration and implemention. declaration: `d5a19e4525/torch/_C/__init__.pyi.in (L157-L162)` Implementation: `d5a19e4525/torch/csrc/Event.cpp (L30-L32)` Question: Which one should we choose? - Change enable_timing to False to be consistent with torch.cuda.Event - Change enable_timing to True to avoid BC-break Pull Request resolved: https://github.com/pytorch/pytorch/pull/151221 Approved by: https://github.com/albanD ghstack dependencies: #151404	2025-04-26 13:51:56 +00:00
Chuanqi Xu	4a46ee96d2	[Indcutor Remote Cache] Raise an exception if redis module is required but not available (#151779 ) If we need redis but redis is not available, it is better to tell the user to install redis instead of continue silently. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151779 Approved by: https://github.com/aorenste	2025-04-26 11:21:54 +00:00
Mu-Chu Lee	8d427e9e76	[AOTInductor] Inherit Buffer if not being updated (#152092 ) Summary: Inherit buffer from original constants buffer if it's not being updated. Test Plan: TBD Differential Revision: D73571260 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152092 Approved by: https://github.com/kflu, https://github.com/jingsh	2025-04-26 04:28:23 +00:00
Dan Johnson	d22c4cc353	Add option to use mempool on OOM (#151487 ) MemPool is a separate pool of memory handled by the caching allocator. This PR adds the option let the caching allocator try to use this pool as a last resort instead of OOMing by associating a use_on_oom bool with each MemPool. Usage: Users can optionally specify a ``use_on_oom`` bool (which is False by default) during MemPool creation. If true, then the CUDACachingAllocator will be able to use memory in this pool as a last resort instead of OOMing. ``` pool = torch.cuda.MemPool(allocator, use_on_oom=True) with torch.cuda.use_mem_pool(pool): a = torch.randn(40 * 1024 * 1024, dtype=torch.uint8, device="cuda") del a # at the memory limit, this will succeed by using pool's memory in order to avoid the oom b = torch.randn(40 * 1024 * 1024, dtype=torch.uint8, device="cuda") ``` Testing: ``` python test/test_cuda.py -k test_mempool_limited_memory_with_allocator ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/151487 Approved by: https://github.com/eqy, https://github.com/syed-ahmed, https://github.com/ngimel	2025-04-26 04:04:57 +00:00
cyy	65b845f82b	Remove useless options for third-party ONNX build (#147616 ) Treat ONNX CMake targets properly and remove unneeded options. Pull Request resolved: https://github.com/pytorch/pytorch/pull/147616 Approved by: https://github.com/malfet	2025-04-26 02:34:08 +00:00
Alexander Grund	d9d306e8e9	Fix inductor test_linear_with_in_out_buffer (#151548 ) Without MKL there is only 1 epilogue, not 2 because `addmm` is used instead of `packed_linear/_mkl_linear`. This fails first at `TestSelectAlgorithmCPU.test_linear_with_in_out_buffer_batch_size_8_in_features_3_in_features2_192_image_size_224_out_features_64_bias_True_cpu_float32` Instead of skipping the whole test just adjust the count for the single check. Final numbers of `test/inductor/test_cpu_select_algorithm.py` without MKL: ``` Ran 1337 tests OK (skipped=1211) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/151548 Approved by: https://github.com/jansel	2025-04-26 01:53:34 +00:00
Michal Gallus	0e015ef116	[ROCm][Windows] Fix HIP Caffe2 Tests (#152014 ) Solves the following problems of caffe2 HIP tests building on Windows: 1. HIP tests now use `hip_add_executable` to be built with custom_command invoking hip compiler, due to lack of cmake support for HIP in 3.18 (currently used). 2. failing with "Command line too long" which resulted from `hip_add_executable` adding the same flags over and over on top of `HIP_HIPCC_FLAGS` with every test added. 3. Disables `HasSameArgTypes` test on Windows, as `at::native::modern::detail` is nowhere to be found in the codebase (I think it must be a legacy thing). Perhaps the whole test should be removed/rewritten? Pull Request resolved: https://github.com/pytorch/pytorch/pull/152014 Approved by: https://github.com/jeffdaily	2025-04-26 01:35:46 +00:00
Nikita Shulga	3ef6d6924a	[BE] Switch `TestConsistency` to MPS device (#147893 ) Which will eventually allow move decorators away more `common_mps.py` Adjust tolerances accordingly. XFAIL a bunch of tests on MacOS-13, which is going to be deprecated anyway Pull Request resolved: https://github.com/pytorch/pytorch/pull/147893 Approved by: https://github.com/atalman ghstack dependencies: #152204	2025-04-26 01:19:21 +00:00
Nikita Shulga	73f11e3365	[BE] Do not allow PyTorch codebase to use `c10::optional` (#150464 ) Extensions can still rely on it, and we should decorate it with deprecated, but it is a C++20 feature. XPU still uses it, so exclude XPU builds until https://github.com/intel/torch-xpu-ops/pull/1615 is merged Test plan: - `0def9b4acc` should fail MPS builds ``` /Users/ec2-user/runner/_work/pytorch/pytorch/aten/src/ATen/native/mps/OperationUtils.mm:975:44: error: no template named 'optional' in namespace 'c10'; did you mean 'std::optional'? c10::optional<int64_t> extra) { ^~~~~~~~~~~~~ std::optional ``` - `a769759dd4` should fail CUDA builds ``` /var/lib/jenkins/workspace/torch/csrc/distributed/c10d/CUDASymmetricMemoryOps.cu(530): error: namespace "c10" has no member "nullopt" input, c10::nullopt, reduce_op, group_name, out); ^ 1 error detected in the compilation of ``` Fixes https://github.com/pytorch/pytorch/issues/150313 Pull Request resolved: https://github.com/pytorch/pytorch/pull/150464 Approved by: https://github.com/atalman	2025-04-26 01:15:53 +00:00
Flavio Sales Truzzi	4647658247	[PT2] - Allowlist should have precedence (#151942 ) Summary: When working on List[List[int]], the ints were being considered Constants regardless of their inclusion on the allowlist. Test Plan: CI + new test https://www.internalfb.com/intern/testinfra/testrun/5066549856504774 Differential Revision: D73137631 Pull Request resolved: https://github.com/pytorch/pytorch/pull/151942 Approved by: https://github.com/laithsakka	2025-04-26 00:58:43 +00:00
PyTorch MergeBot	fa1b4ef649	Revert "Rewrite the guts of torch::jit::Lexer to speed it up (#151850 )" This reverts commit `47d34261e0`. Reverted https://github.com/pytorch/pytorch/pull/151850 on behalf of https://github.com/ZainRizvi due to This codev PR is breaking on it's internal counterpart diff D73129443. For codev PRs like this one, please always make sure the internal diff is green and then land the diff internally. The Github PR will be automatically merged ([comment](https://github.com/pytorch/pytorch/pull/151850#issuecomment-2831686141))	2025-04-26 00:44:11 +00:00
Scott Wolchok	47d34261e0	Rewrite the guts of torch::jit::Lexer to speed it up (#151850 ) The trie-based approach was, apparently, not efficient. This incidentally fixes a bug where "not inp" and "is note" were lexed incorrectly; see test_lexer.cpp update. Differential Revision: [D73129443](https://our.internmc.facebook.com/intern/diff/D73129443/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/151850 Approved by: https://github.com/Skylion007 ghstack dependencies: #151801, #151802, #151803, #151804, #151805, #151806, #151807, #151810, #151849	2025-04-25 23:49:35 +00:00
PyTorch MergeBot	0f765773e3	Revert "[BE] Do not allow PyTorch codebase to use `c10::optional` (#150464 )" This reverts commit `490ef768cf`. Reverted https://github.com/pytorch/pytorch/pull/150464 on behalf of https://github.com/clee2000 due to broke xpu [GH job link](https://github.com/pytorch/pytorch/actions/runs/14674243034/job/41187443432) [HUD commit link](`490ef768cf`)? ([comment](https://github.com/pytorch/pytorch/pull/150464#issuecomment-2831608162))	2025-04-25 23:34:56 +00:00
Chien-Chin Huang	6aa92806db	[CP] Use TorchFunctionMode to dispatch SDPA for CP (#147902 ) While we prefer not use monkey patching to dispatch SDPA, TorchFunctionMode is currently not compatible with selective activation checkpointing (https://github.com/pytorch/pytorch/issues/147995). This PR adds `TorchFunctionMode` to CP code and make it configurable. Pull Request resolved: https://github.com/pytorch/pytorch/pull/147902 Approved by: https://github.com/XilunWu	2025-04-25 23:33:48 +00:00
Davide Italiano	e28864fc0f	[MPS/inductor] Fix the approximation of polygamma for n == 0. (#152214 ) Fixes #152205 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152214 Approved by: https://github.com/malfet	2025-04-25 22:42:45 +00:00
Scott Wolchok	cf101d66ee	Add simple direct C++ tests for torch::jit::Lexer (#151849 ) We have test_jit.py, but given that I'm working on significant changes to the lexer, it seems nice to have direct C++ tests. (Also, writing the tests caught a pair of related bugs; see the two tests with "Bug" in their name. The rewrite will fix them.) Differential Revision: [D73402367](https://our.internmc.facebook.com/intern/diff/D73402367/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/151849 Approved by: https://github.com/malfet ghstack dependencies: #151801, #151802, #151803, #151804, #151805, #151806, #151807, #151810	2025-04-25 22:39:49 +00:00
Nikita Shulga	490ef768cf	[BE] Do not allow PyTorch codebase to use `c10::optional` (#150464 ) Extensions can still rely on it, and we should decorate it with deprecated, but it is a C++20 feature Test plan: - `0def9b4acc` should fail MPS builds ``` /Users/ec2-user/runner/_work/pytorch/pytorch/aten/src/ATen/native/mps/OperationUtils.mm:975:44: error: no template named 'optional' in namespace 'c10'; did you mean 'std::optional'? c10::optional<int64_t> extra) { ^~~~~~~~~~~~~ std::optional ``` - `a769759dd4` should fail CUDA builds ``` /var/lib/jenkins/workspace/torch/csrc/distributed/c10d/CUDASymmetricMemoryOps.cu(530): error: namespace "c10" has no member "nullopt" input, c10::nullopt, reduce_op, group_name, out); ^ 1 error detected in the compilation of ``` Fixes https://github.com/pytorch/pytorch/issues/150313 Pull Request resolved: https://github.com/pytorch/pytorch/pull/150464 Approved by: https://github.com/atalman	2025-04-25 22:03:48 +00:00
Anthony Shoumikhin	9e50c21e27	Fix xrefs (#151888 ) Fix existing cross references and removed old ones Pull Request resolved: https://github.com/pytorch/pytorch/pull/151888 Approved by: https://github.com/eqy, https://github.com/huydhn, https://github.com/svekars	2025-04-25 21:27:27 +00:00

1 2 3 4 5 ...

87023 Commits