Skip to content

Use relocatable IR linked at load time#3200

Open
maleadt wants to merge 9 commits into
mainfrom
tb/relocations
Open

Use relocatable IR linked at load time#3200
maleadt wants to merge 9 commits into
mainfrom
tb/relocations

Conversation

@maleadt

@maleadt maleadt commented Jul 14, 2026

Copy link
Copy Markdown
Member

@maleadt
maleadt marked this pull request as draft July 14, 2026 13:35
@maleadt
maleadt marked this pull request as ready for review July 15, 2026 20:34
@github-actions

github-actions Bot commented Jul 16, 2026

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: 7d26d61 Previous: 069cdff Ratio
array/accumulate/Float32/1d 97592 ns 98074 ns 1.00
array/accumulate/Float32/dims=1 71237 ns 74919 ns 0.95
array/accumulate/Float32/dims=1L 1599227 ns 1600002 ns 1.00
array/accumulate/Float32/dims=2 136868 ns 140367 ns 0.98
array/accumulate/Float32/dims=2L 659994 ns 659745 ns 1.00
array/accumulate/Int64/1d 117120 ns 117818 ns 0.99
array/accumulate/Int64/dims=1 75019 ns 79251 ns 0.95
array/accumulate/Int64/dims=1L 1715771 ns 1717606 ns 1.00
array/accumulate/Int64/dims=2 147864 ns 153507 ns 0.96
array/accumulate/Int64/dims=2L 986380 ns 986800 ns 1.00
array/broadcast 17598 ns 18060 ns 0.97
array/broadcast launch 8519 ns
array/construct 836.2714285714286 ns 869.1052631578947 ns 0.96
array/copy 16095 ns 15967 ns 1.01
array/copyto!/cpu_to_gpu 209326 ns 208599 ns 1.00
array/copyto!/gpu_to_cpu 240849 ns 241280 ns 1.00
array/copyto!/gpu_to_gpu 8881.333333333334 ns 9209.333333333334 ns 0.96
array/iteration/findall/bool 131187 ns 132021 ns 0.99
array/iteration/findall/int 144614 ns 145286 ns 1.00
array/iteration/findfirst/bool 67629 ns 67274 ns 1.01
array/iteration/findfirst/int 69032 ns 68529 ns 1.01
array/iteration/findmin/1d 61297 ns 64309 ns 0.95
array/iteration/findmin/2d 98475 ns 99940 ns 0.99
array/iteration/logical 182849 ns 185815 ns 0.98
array/iteration/scalar 60408 ns 63136 ns 0.96
array/permutedims/2d 47743 ns 48431 ns 0.99
array/permutedims/3d 48903 ns 50026 ns 0.98
array/permutedims/4d 49377 ns 49849 ns 0.99
array/random/rand/Float32 10879 ns 11669 ns 0.93
array/random/rand/Int64 20061 ns 22883 ns 0.88
array/random/rand!/Float32 7772.25 ns 7731.75 ns 1.01
array/random/rand!/Int64 19199 ns 20210 ns 0.95
array/random/randn/Float32 32449 ns 32574 ns 1.00
array/random/randn!/Float32 23327 ns 23189 ns 1.01
array/reductions/mapreduce/Float32/1d 31326 ns 31371 ns 1.00
array/reductions/mapreduce/Float32/dims=1 36979 ns 37491 ns 0.99
array/reductions/mapreduce/Float32/dims=1L 49742 ns 49977 ns 1.00
array/reductions/mapreduce/Float32/dims=2 54349 ns 54508 ns 1.00
array/reductions/mapreduce/Float32/dims=2L 66057 ns 66329 ns 1.00
array/reductions/mapreduce/Int64/1d 38488 ns 37643 ns 1.02
array/reductions/mapreduce/Int64/dims=1 40321 ns 39961 ns 1.01
array/reductions/mapreduce/Int64/dims=1L 87930 ns 87917 ns 1.00
array/reductions/mapreduce/Int64/dims=2 56932 ns 57047 ns 1.00
array/reductions/mapreduce/Int64/dims=2L 82277 ns 82751 ns 0.99
array/reductions/reduce/Float32/1d 31846 ns 31482 ns 1.01
array/reductions/reduce/Float32/dims=1 37376 ns 37630 ns 0.99
array/reductions/reduce/Float32/dims=1L 50055 ns 50073 ns 1.00
array/reductions/reduce/Float32/dims=2 54492 ns 54518 ns 1.00
array/reductions/reduce/Float32/dims=2L 67439 ns 67963 ns 0.99
array/reductions/reduce/Int64/1d 38401 ns 38054 ns 1.01
array/reductions/reduce/Int64/dims=1 39983 ns 39983 ns 1
array/reductions/reduce/Int64/dims=1L 87863 ns 87988 ns 1.00
array/reductions/reduce/Int64/dims=2 56796 ns 56860 ns 1.00
array/reductions/reduce/Int64/dims=2L 82392 ns 82191 ns 1.00
array/reverse/1d 16230 ns 16098 ns 1.01
array/reverse/1dL 68827 ns 68844 ns 1.00
array/reverse/1dL_inplace 67029 ns 66918 ns 1.00
array/reverse/1d_inplace 8374 ns 9757 ns 0.86
array/reverse/2d 19561 ns 19108 ns 1.02
array/reverse/2dL 72876 ns 72561 ns 1.00
array/reverse/2dL_inplace 66835 ns 66785 ns 1.00
array/reverse/2d_inplace 9480 ns 9397 ns 1.01
array/sorting/1d 2639323 ns 2656939 ns 0.99
array/sorting/2d 1027902 ns 1038211 ns 0.99
array/sorting/by 3182088 ns 3192363 ns 1.00
cuda/synchronization/context/auto 1041.8 ns 1015.3 ns 1.03
cuda/synchronization/context/blocking 801.010752688172 ns 784.6960784313726 ns 1.02
cuda/synchronization/context/nonblocking 5713 ns 5655.166666666667 ns 1.01
cuda/synchronization/stream/auto 877.8775510204082 ns 857.3174603174604 ns 1.02
cuda/synchronization/stream/blocking 679.5490196078431 ns 654.7730061349694 ns 1.04
cuda/synchronization/stream/nonblocking 5532.285714285715 ns 5491 ns 1.01
integration/byval/reference 147329 ns 147245 ns 1.00
integration/byval/slices=1 151637 ns 149267 ns 1.02
integration/byval/slices=2 294507 ns 292191 ns 1.01
integration/byval/slices=3 437547 ns 434868 ns 1.01
integration/cudadevrt 104428 ns 104357 ns 1.00
integration/volumerhs 9141010 ns 9315717 ns 0.98
kernel/indexing 12751 ns 12183 ns 1.05
kernel/indexing_checked 13528 ns 13078 ns 1.03
kernel/launch 2173.8888888888887 ns 2032.3333333333333 ns 1.07
kernel/occupancy 685.8343949044586 ns 667.8466257668712 ns 1.03
kernel/rand 13777 ns 14903 ns 0.92
latency/import 4092841703 ns 4116528672 ns 0.99
latency/precompile 4782383781 ns 4814427609 ns 0.99
latency/ttfp 4955136012 ns 5198634459 ns 0.95

This comment was automatically generated by workflow using github-action-benchmark.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant