pytorch

mirror of https://github.com/zebrajr/pytorch.git synced 2025-12-06 12:20:52 +01:00

History

xiaobing.zhang b47e9b97a2 Add op bitwise_and (#31104 ) Summary: Refer to https://github.com/pytorch/pytorch/pull/25665, add `bitwise_and` operator. Benchmark script : ``` import timeit #for __and__ for n, t in [(10, 100000),(1000, 10000)]: print('__and__ (a.numel() == {}) for {} times'.format(n, t)) for device in ('cpu', 'cuda'): for dtype in ('torch.int8', 'torch.uint8', 'torch.int16', 'torch.int32', 'torch.int64'): print(f'device: {device}, dtype: {dtype}, {t} times', end='\t\t') print(timeit.timeit(f'a & b\nif "{device}" == "cuda": torch.cuda.synchronize()', setup=f'import torch; a = torch.randint(0, 10, ({n},), dtype = {dtype}, device="{device}"); b = torch.randint(0, 10, ({n},), dtype = {dtype}, device="{device}")', number=t)) #for __iand__ for n, t in [(10, 100000),(1000, 10000)]: print('__iand__ (a.numel() == {}) for {} times'.format(n, t)) for device in ('cpu', 'cuda'): for dtype in ('torch.int8', 'torch.uint8', 'torch.int16', 'torch.int32', 'torch.int64'): print(f'device: {device}, dtype: {dtype}, {t} times', end='\t\t') print(timeit.timeit(f'a & b\nif "{device}" == "cuda": torch.cuda.synchronize()', setup=f'import torch; a = torch.randint(0, 10, ({n},), dtype = {dtype}, device="{device}"); b = torch.tensor(5, dtype = {dtype}, device="{device}")', number=t)) ``` Device: Tesla P100, skx-8180 Cuda verison: 9.0.176 Before: ``` __and__ (a.numel() == 10) for 100000 times device: cpu, dtype: torch.int8, 100000 times 0.1766007635742426 device: cpu, dtype: torch.uint8, 100000 times 0.17322628945112228 device: cpu, dtype: torch.int16, 100000 times 0.17650844901800156 device: cpu, dtype: torch.int32, 100000 times 0.17711848113685846 device: cpu, dtype: torch.int64, 100000 times 0.18240160401910543 device: cuda, dtype: torch.int8, 100000 times 1.273967768996954 device: cuda, dtype: torch.uint8, 100000 times 1.2778537990525365 device: cuda, dtype: torch.int16, 100000 times 1.2753686187788844 device: cuda, dtype: torch.int32, 100000 times 1.2797665279358625 device: cuda, dtype: torch.int64, 100000 times 1.2933144550770521 __and__ (a.numel() == 1000) for 10000 times device: cpu, dtype: torch.int8, 10000 times 0.031139614060521126 device: cpu, dtype: torch.uint8, 10000 times 0.03091452084481716 device: cpu, dtype: torch.int16, 10000 times 0.022756479680538177 device: cpu, dtype: torch.int32, 10000 times 0.025045674294233322 device: cpu, dtype: torch.int64, 10000 times 0.024164282716810703 device: cuda, dtype: torch.int8, 10000 times 0.12820732593536377 device: cuda, dtype: torch.uint8, 10000 times 0.12775669433176517 device: cuda, dtype: torch.int16, 10000 times 0.12697868794202805 device: cuda, dtype: torch.int32, 10000 times 0.12832533661276102 device: cuda, dtype: torch.int64, 10000 times 0.1280576130375266 __iand__ (a.numel() == 10) for 100000 times device: cpu, dtype: torch.int8, 100000 times 0.3687064303085208 device: cpu, dtype: torch.uint8, 100000 times 0.36253443732857704 device: cpu, dtype: torch.int16, 100000 times 0.362891579978168 device: cpu, dtype: torch.int32, 100000 times 0.37680106051266193 device: cpu, dtype: torch.int64, 100000 times 0.3689364707097411 device: cuda, dtype: torch.int8, 100000 times 1.419940729625523 device: cuda, dtype: torch.uint8, 100000 times 1.4247053815051913 device: cuda, dtype: torch.int16, 100000 times 1.4191444097086787 device: cuda, dtype: torch.int32, 100000 times 1.4305962566286325 device: cuda, dtype: torch.int64, 100000 times 1.4567416654899716 __iand__ (a.numel() == 1000) for 10000 times device: cpu, dtype: torch.int8, 10000 times 0.06224383972585201 device: cpu, dtype: torch.uint8, 10000 times 0.06205617543309927 device: cpu, dtype: torch.int16, 10000 times 0.05016433447599411 device: cpu, dtype: torch.int32, 10000 times 0.05216377507895231 device: cpu, dtype: torch.int64, 10000 times 0.06139362137764692 device: cuda, dtype: torch.int8, 10000 times 0.14827249851077795 device: cuda, dtype: torch.uint8, 10000 times 0.14801877550780773 device: cuda, dtype: torch.int16, 10000 times 0.14952312968671322 device: cuda, dtype: torch.int32, 10000 times 0.14999118447303772 device: cuda, dtype: torch.int64, 10000 times 0.14951884001493454 ``` After: ``` __and__ (a.numel() == 10) for 100000 times device: cpu, dtype: torch.int8, 100000 times 0.23157884553074837 device: cpu, dtype: torch.uint8, 100000 times 0.23063660878688097 device: cpu, dtype: torch.int16, 100000 times 0.23005440644919872 device: cpu, dtype: torch.int32, 100000 times 0.23748818412423134 device: cpu, dtype: torch.int64, 100000 times 0.24106105230748653 device: cuda, dtype: torch.int8, 100000 times 1.4394256137311459 device: cuda, dtype: torch.uint8, 100000 times 1.4436759827658534 device: cuda, dtype: torch.int16, 100000 times 1.4631587155163288 device: cuda, dtype: torch.int32, 100000 times 1.459101552143693 device: cuda, dtype: torch.int64, 100000 times 1.4784048134461045 __and__ (a.numel() == 1000) for 10000 times device: cpu, dtype: torch.int8, 10000 times 0.028442862443625927 device: cpu, dtype: torch.uint8, 10000 times 0.028130197897553444 device: cpu, dtype: torch.int16, 10000 times 0.025318274274468422 device: cpu, dtype: torch.int32, 10000 times 0.02519288007169962 device: cpu, dtype: torch.int64, 10000 times 0.028299466706812382 device: cuda, dtype: torch.int8, 10000 times 0.14342594426125288 device: cuda, dtype: torch.uint8, 10000 times 0.145280827768147 device: cuda, dtype: torch.int16, 10000 times 0.14673697855323553 device: cuda, dtype: torch.int32, 10000 times 0.14499565307050943 device: cuda, dtype: torch.int64, 10000 times 0.14582364354282618 __iand__ (a.numel() == 10) for 100000 times device: cpu, dtype: torch.int8, 100000 times 0.25548241566866636 device: cpu, dtype: torch.uint8, 100000 times 0.2552562616765499 device: cpu, dtype: torch.int16, 100000 times 0.25905191246420145 device: cpu, dtype: torch.int32, 100000 times 0.26635489892214537 device: cpu, dtype: torch.int64, 100000 times 0.26269810926169157 device: cuda, dtype: torch.int8, 100000 times 1.485458506271243 device: cuda, dtype: torch.uint8, 100000 times 1.4742380809038877 device: cuda, dtype: torch.int16, 100000 times 1.507783885113895 device: cuda, dtype: torch.int32, 100000 times 1.4926990242674947 device: cuda, dtype: torch.int64, 100000 times 1.519851053133607 __iand__ (a.numel() == 1000) for 10000 times device: cpu, dtype: torch.int8, 10000 times 0.03425929415971041 device: cpu, dtype: torch.uint8, 10000 times 0.03293587639927864 device: cpu, dtype: torch.int16, 10000 times 0.029559112153947353 device: cpu, dtype: torch.int32, 10000 times 0.030915481969714165 device: cpu, dtype: torch.int64, 10000 times 0.03292469773441553 device: cuda, dtype: torch.int8, 10000 times 0.15792148280888796 device: cuda, dtype: torch.uint8, 10000 times 0.16000914946198463 device: cuda, dtype: torch.int16, 10000 times 0.1600684942677617 device: cuda, dtype: torch.int32, 10000 times 0.16162546630948782 device: cuda, dtype: torch.int64, 10000 times 0.1629159888252616 ``` Fix https://github.com/pytorch/pytorch/issues/24508, https://github.com/pytorch/pytorch/issues/24509, https://github.com/pytorch/pytorch/issues/24655, https://github.com/pytorch/pytorch/issues/24656. Pull Request resolved: https://github.com/pytorch/pytorch/pull/31104 Differential Revision: D18938930 Pulled By: VitalyFedyunin fbshipit-source-id: a77e805a0b84e8ace16c6e648c2f67dad44f2e44		2020-01-03 10:32:36 -08:00
..
_static	Improve documentation around builtin functions (#30347 )	2019-12-04 13:50:40 -08:00
_templates	Generate sphinx docs with secure content. (#18508 )	2019-03-27 11:01:48 -07:00
community	Update persons_of_interest.rst	2019-12-05 21:20:40 -08:00
notes	Add the __torch_function__ API override mechanism (#30730 )	2019-12-04 13:19:07 -08:00
org/pytorch	Revert D17850696: [pytorch][PR] Updates to quantization related files, index.rst, and javadocs	2019-10-10 09:23:33 -07:00
scripts	Add torch.nn.GELU for GELU activation (#28944 )	2019-11-03 21:55:05 -08:00
__config__.rst	Allow a non-OpenMP based build (#19749 )	2019-05-06 19:34:48 -07:00
autograd.rst	Added docs for context method mixins. Fixes issue #27365 (#28643 )	2019-10-28 08:31:35 -07:00
bottleneck.rst	[docs] Clarify more CUDA profiling gotchas in bottleneck docs (#6763 )	2018-04-19 13:15:27 -04:00
checkpoint.rst	Stashing checkpointing RNG states based on devices of arg tensors (#14518 )	2018-12-11 09:48:45 -08:00
conf.py	Improve documentation around builtin functions (#30347 )	2019-12-04 13:50:40 -08:00
cpp_extension.rst	Inline JIT C++ Extensions (#7059 )	2018-04-30 11:48:44 -04:00
cuda_deterministic_backward.rst	Typo correction in cuda_deterministic_backward.rst (#25011 )	2019-08-22 21:19:39 -07:00
cuda_deterministic.rst	Amend nondeterminism notes (#12217 )	2018-10-16 23:59:26 -07:00
cuda.rst	Fix most documentation warnings (#27782 )	2019-10-13 10:34:01 -07:00
cudnn_deterministic.rst	Amend nondeterminism notes (#12217 )	2018-10-16 23:59:26 -07:00
cudnn_persistent_rnn.rst	don't copy weight gradients in rnn (#12600 )	2018-10-12 13:34:10 -07:00
data.rst	Fix typo in data.rst docs	2019-12-18 09:52:10 -08:00
distributed.rst	Fix typos (#30606 )	2019-12-02 20:17:42 -08:00
distributions.rst	Revert D18249048: Moved VonMises distribution with sampling upstream from Pyro.	2019-11-04 09:50:50 -08:00
dlpack.rst	document torch.utils.dlpack (#9343 )	2018-07-11 07:46:09 -07:00
hub.rst	Fix typos (#30606 )	2019-12-02 20:17:42 -08:00
index.rst	Add docs for RPC, dist autograd, and RRef modules (#29276 )	2019-11-14 14:32:03 -08:00
jit_builtin_functions.rst	Fix builtin function reference (#24056 )	2019-08-09 15:58:15 -07:00
jit_language_reference.rst	Cleanup after moving language reference (#31146 )	2019-12-18 15:09:35 -08:00
jit_python_reference.rst	Add Python language reference docs (#30686 )	2019-12-26 13:21:36 -08:00
jit_unsupported.rst	add unsupported section (#31329 )	2019-12-18 13:56:02 -08:00
jit.rst	Add Python language reference docs (#30686 )	2019-12-26 13:21:36 -08:00
math-quantizer-equation.png	adding quantization.rst file for quantization feature (#27559 )	2019-10-09 16:45:09 -07:00
model_zoo.rst	add/move a few apis in torch.hub (#18758 )	2019-04-10 23:10:39 -07:00
multiprocessing.rst	Bag of documentation fixes; fix more sphinx warnings (#27850 )	2019-10-15 07:31:14 -07:00
name_inference.rst	Fix typos (#30606 )	2019-12-02 20:17:42 -08:00
named_tensor.rst	Bag of documentation fixes; fix more sphinx warnings (#27850 )	2019-10-15 07:31:14 -07:00
nn.functional.rst	Breaks up NN module in docs so it loads faster.	2019-06-11 09:38:41 -07:00
nn.init.rst	Bag of documentation fixes; fix more sphinx warnings (#27850 )	2019-10-15 07:31:14 -07:00
nn.rst	Pruning Functionality (#24076 )	2019-11-08 19:38:00 -08:00
onnx.rst	Update onnx landing page for 1.3 (#27581 )	2019-10-11 20:53:50 -07:00
optim.rst	Fix capitalization inconsistency in optim.rst	2019-12-04 08:17:03 -08:00
packages.rst	Revert D17850696: [pytorch][PR] Updates to quantization related files, index.rst, and javadocs	2019-10-10 09:23:33 -07:00
quantization.rst	Updates to quantization documentation (#30288 )	2019-11-23 09:29:30 -08:00
random.rst	Fix most documentation warnings (#27782 )	2019-10-13 10:34:01 -07:00
rpc.rst	Document WorkerInfo and RpcBackendOptions structures in RPC docs. (#31077 )	2019-12-11 11:39:57 -08:00
sparse.rst	Bag of documentation fixes; fix more sphinx warnings (#27850 )	2019-10-15 07:31:14 -07:00
storage.rst	Start documenting torch.Tensor (#377 )	2016-12-30 01:21:34 -05:00
tensor_attributes.rst	Expose a torch.result_type and simplify tensor iterator	2019-09-25 06:52:23 -07:00
tensorboard.rst	Add method add_hparams to API doc (#27344 )	2019-10-03 17:07:45 -07:00
tensors.rst	Add op bitwise_and (#31104 )	2020-01-03 10:32:36 -08:00
torch.rst	Add op bitwise_and (#31104 )	2020-01-03 10:32:36 -08:00
type_info.rst	Allow converting char tensor to numpy; add [fi]info.min (#15046 )	2018-12-24 09:11:24 -08:00