Retrieved article excerpt
Open article · Retrieved 2026-10-07T14:27:41.467364+00:00
# [Bringing SYCL to Android: A Vulkan backend for portable GPU compute](https://adaptivecpp.github.io/hipsycl/adaptivecpp/sycl/vulkan/android/vulkan-android/)
13 minute read
# Introduction
SYCL’s promise is performance portability: write modern C++ once and execute across many different accelerators. But that
promise only goes as far as the available backends. While desktop and HPC platforms capitalize on established OpenCL,
CUDA, HIP, and Level Zero backends for accelerating SYCL applications on the GPU, mobile isn’t a domain commonly associated
with SYCL. Yet almost every Android device ships with a capable Vulkan implementation, making mobile an unfulfilled chapter of
the SYCL performance portability story.
While projects such as [Sylkan](https://dl.acm.org/doi/10.1145/3456669.3456683) have demonstrated that SYCL over Vulkan is
feasible, the space has remained relatively unexplored in terms of feature completeness, Android support, and integration
with the wider SYCL ecosystem. OpenCL has had more success in this area, with projects such as
[clvk](https://github.com/kpet/clvk), [pocl](https://github.com/pocl/pocl), and [ANGLE](https://github.com/google/angle)
successfully layering the OpenCL compute API over Vulkan.
AdaptiveCpp closes this SYCL gap with its new Vulkan backend and Android cross-compilation support.
Here is what we’ll cover in this blog post:
- [Getting Started](https://adaptivecpp.github.io/hipsycl/adaptivecpp/sycl/vulkan/android/vulkan-android/#getting-started): How to build AdaptiveCpp with the Vulkan backend on desktop so you can try it yourself.
- [Benchmarks](https://adaptivecpp.github.io/hipsycl/adaptivecpp/sycl/vulkan/android/vulkan-android/#android-benchmarks): Early performance results from bringing SYCL GPU acceleration to Android.
- [Under the Hood](https://adaptivecpp.github.io/hipsycl/adaptivecpp/sycl/vulkan/android/vulkan-android/#under-the-hood): A deep dive into how AdaptiveCpp layers over Vulkan, exploring the runtime and the compiler.
# Getting Started
When it comes to using the Vulkan AdaptiveCpp backend, thanks to the proliferation of Vulkan drivers,
there are many platforms on which the backend can be tested. We’re proud that AdaptiveCpp GitHub CI now has
all of Linux, Windows, and macOS operating systems tested on the Vulkan backend on every commit,
using [Mesa llvmpipe](https://docs.mesa3d.org/drivers/llvmpipe.html) for Linux and Windows,
and [MoltenVK](https://github.com/KhronosGroup/MoltenVK) for macOS.
In this article, we’ll only cover how to build and use the backend on Ubuntu using a native build flow.
Android requires a more complex build process using the Android Native Development Kit (NDK) to
cross-compile AdaptiveCpp and other dependencies. You can find the in-depth instructions for how to do that
[here](https://github.com/AdaptiveCpp/AdaptiveCpp/blob/develop/doc/install-android.md).
## Building AdaptiveCpp With The Vulkan Backend
There are four main pieces to assemble before you can compile and run your first SYCL program for Vulkan: the [LunarG SDK](https://adaptivecpp.github.io/hipsycl/adaptivecpp/sycl/vulkan/android/vulkan-android/#lunarg-sdk),
a [Vulkan driver](https://adaptivecpp.github.io/hipsycl/adaptivecpp/sycl/vulkan/android/vulkan-android/#vulkan-driver), the [clspv compiler](https://adaptivecpp.github.io/hipsycl/adaptivecpp/sycl/vulkan/android/vulkan-android/#clspv), and [AdaptiveCpp](https://adaptivecpp.github.io/hipsycl/adaptivecpp/sycl/vulkan/android/vulkan-android/#adaptivecpp) itself.
The first build dependency is the LunarG Vulkan SDK which contains SPIR-V Tools, Vulkan layers and the loader, and the VulkanHpp headers.
The second is a `clspv` executable, which can be built following the GitHub repo instructions. A Vulkan driver is also required to run the backend.
Finally, we need to build AdaptiveCpp with the Vulkan backend enabled.
### LunarG SDK
Download the Vulkan SDK tarball from the [LunarG website](https://vulkan.lunarg.com/sdk/home) and decompress it.
The LunarG SDK can then be made available in your system path after sourcing the `setup-env.sh` script that it
ships (we recommend automatically sourcing this as part of your `.bashrc`).
```
$ wget https://sdk.lunarg.com/sdk/download/1.4.357.0/linux/vulkansdk-linux-x86_64-1.4.357.0.tar.xz
$ tar -xvf vulkansdk-linux-x86_64-1.4.357.0.tar.xz
$ source 1.4.357.0/setup-env.sh
```
### Vulkan Driver
The Vulkan SDK comes with a `vulkaninfo` tool for printing the Vulkan drivers on your system.
At least one driver is required to use as a SYCL backend device. If you don’t have any installed then the
easiest way to reliably get a supported driver is to install the Mesa drivers with
`apt install mesa-vulkan-drivers`. This will provide at least the llvmpipe CPU Vulkan driver
which provides all the necessary capabilities for SYCL. For example:
```
$ vulkaninfo --summary
GPU0:
apiVersion = 1.4.318
driverVersion = 25.2.8
vendorID = 0x10005
deviceID = 0x0000
deviceType = PHYSICAL_DEVICE_TYPE_CPU
deviceName = llvmpipe (LLVM 20.1.8, 256 bits)
driverID = DRIVER_ID_MESA_LLVMPIPE
driverName = llvmpipe
driverInfo = Mesa 25.2.8-0ubuntu0.25.10.2 (LLVM 20.1.8)
conformanceVersion = 1.3.1.1
```
### clspv
A `clspv` executable is required to be invoked at runtime as part of runtime kernel compilation,
and can be built from [source](https://github.com/google/clspv). The exact commit of clspv should be checked in AdaptiveCpp CI
or [doc/install-vulkan.md](https://github.com/AdaptiveCpp/AdaptiveCpp/blob/develop/doc/install-vulkan.md#requirements)
for the SHA hash.
```
$ git clone https://github.com/google/clspv
$ cd clspv
$ git checkout <supported commit>
$ python3 utils/fetch_sources.py
$ mkdir build && cd build
$ cmake .. -GNinja
$ ninja
$ export CLSPV_BIN_DIR=$PWD/bin
```
The path to the directory with the `clspv` tool is exported as an environment variable so we
can reference it in a later build step.
### AdaptiveCpp
Now we can build AdaptiveCpp itself using these dependencies. If you haven’t got it already,
you first need to clone the [AdaptiveCpp GitHub repo](https://github.com/AdaptiveCpp/AdaptiveCpp).
Combined with the `-DWITH_VULKAN_BACKEND=ON` option for enabling the Vulkan backend in the build,
the relevant parts of the CMake invocation are:
```
$ git clone https://github.com/AdaptiveCpp/AdaptiveCpp.git
$ cd AdaptiveCpp && mkdir build && cd build
$ cmake -GNinja -DWITH_VULKAN_BACKEND=ON -DCMAKE_PROGRAM_PATH=$CLSPV_BIN_DIR -DCMAKE_INSTALL_PREFIX=$PWD/install
$ ninja install
$ export ACPP_BIN_DIR=$PWD/install/bin
```
You can then check that a Vulkan device is indeed available. Note that other devices may be available too,
but the below is the minimum expected number of devices that `acpp-info` should output.
```
$ $ACPP_BIN_DIR/acpp-info -l
=================Backend information===================
Loaded backend 0: OpenMP
Found device: AdaptiveCpp OpenMP host device
Loaded backend 1: Vulkan
Found device: llvmpipe (LLVM 20.1.8, 256 bits)
```
## Running Applications On The Vulkan Backend
Let’s build and run a simple SYCL application with AdaptiveCpp to show
the Vulkan backend in action. All this application does is use a 1D kernel
to initialize a device USM allocation with the index value of each element,
then verifies on the host that each element was initialized with its expected value.
```
// sycl_test.cpp
#include <iostream>
#include <vector>
#include <sycl/sycl.hpp>
int main() {
sycl::device d{sycl::default_selector{}};
sycl::queue q(d, sycl::property::queue::in_order());
std::string device = d.get_info<sycl::info::device::name>();
std::cout << "Default-selected queue runs on device: " << device << std::endl;
constexpr size_t N = 1024;
int *devicePtr = sycl::malloc_device<int>(N, q);
q.parallel_for(N, [=](sycl::id<1> idx) {
devicePtr[idx] = idx;
});
std::vector<int> dataHost(N);
q.copy(devicePtr, dataHost.data(), N).wait();
bool success = true;
for (int i = 0; i < N; i++) {
success = success && (dataHost[i] == i);
}
std::cout << (success ? "SYCL application SUCCESS" : "SYCL application FAILED")
<< std::endl;
sycl::free(devicePtr, q);
return 0;
}
```
Once we have our compiled `sycl_test` binary we can run it using the `ACPP_VISIBILITY_MASK`
environment variable to select the Vulkan backend as `ACPP_VISIBILITY_MASK=vk`. If there is more
than one compatible Vulkan driver on your system then you can ask for a specific device
by name. For example, here we ask for llvmpipe with `ACPP_VISIBILITY_MASK=vk:llvmpipe`.
```
$ $ACPP_BIN_DIR/acpp sycl_test.cpp -o sycl_test
$ ACPP_VISIBILITY_MASK=vk:llvmpipe ./sycl_test
Default-selected queue runs on device: llvmpipe (LLVM 20.1.8, 256 bits)
SYCL application SUCCESS
```
# Android Benchmarks
With the Ubuntu setup working, let’s look at what this enables on Android by measuring
the benefits of SYCL acceleration on one of the benchmarks from
[HeCBench](https://github.com/ORNL/HeCBench). HeCBench provides multiple source code variants of
each benchmark for different backends. We used the mandelbrot benchmark which has, among others, an OpenMP variant
[mandelbrot-omp](https://github.com/ORNL/HeCBench/tree/master/src/mandelbrot-omp), and a SYCL variant
[mandelbrot-sycl](https://github.com/ORNL/HeCBench/tree/master/src/mandelbrot-sycl).
The benchmarks were cross-compiled using release 27 of the Android NDK. The direct
OpenMP benchmark was compiled as follows:
```
$ cd mandelbrot-omp
$ $NDK/toolchains/llvm/prebuilt/linux-x86_64/bin/clang++ *.cpp -O3 -o mandelbrot-omp-ndk --target=aarch64-linux-android34 -fopenmp=libomp --rtlib=compiler-rt -static-libstdc++
```
Note that the SYCL benchmarks in HeCBench require `-DUSE_GPU=1` to be set during compilation to enable a GPU
SYCL queue selector, so for each benchmark we create two SYCL executables linked against an Android cross-compiled build of AdaptiveCpp
`$ACPP_NDK_BUILD`. See the
[install-android](https://github.com/AdaptiveCpp/AdaptiveCpp/blob/develop/doc/install-android.md)
doc for more details on how to achieve this.
```
$ cd mandelbrot-sycl
$ $ACPP_BIN_DIR/acpp *.cpp -O3 -o mandelbrot-gpu-ndk -DUSE_GPU=1 --target=aarch64-linux-android34 --sysroot=$NDK/toolchains/llvm/prebuilt/linux-x86_64/sysroot --rtlib=compiler-rt -static-libstdc++ -resource-dir=$NDK/toolchains/llvm/prebuilt/linux-x86_64/lib/clang/18/ -L $ACPP_NDK_BUILD/lib
$ $ACPP_BIN_DIR/acpp *.cpp -O3 -o mandelbrot-cpu-ndk --target=aarch64-linux-android34 --sysroot=$NDK/toolchains/llvm/prebuilt/linux-x86_64/sysroot --rtlib=compiler-rt -static-libstdc++ -resource-dir=$NDK/toolchains/llvm/prebuilt/linux-x86_64/lib/clang/18/ -L $ACPP_NDK_BUILD/lib
```
The SYCL acceleration results show the merit of GPU offload. On an Android 14 device with a Qualcomm Snapdragon 8 Gen 3 SoC
containing an Adreno 750 GPU and Arm v8a CPU, we achieved a 46.5% speedup over raw OpenMP by using the GPU exposed by Vulkan.
Using SYCL with the OpenMP backend showed a 7% overhead over the raw OpenMP equivalent, which matches expectations
from the same comparison on a desktop machine.
Taking the average parallel time over 1000 iterations, which is output by the benchmark, we observed the following:
| Benchmark | Average parallel time (ms) |
| --- | --- |
| `./mandelbrot-omp-ndk 1000` | 101 |
| `./mandelbrot-cpu-ndk 1000` | 108 |
| `./mandelbrot-gpu-ndk 1000` | 54 |
*Data taken from the median of 5 runs based on a build with
AdaptiveCpp commit [8a75ba2410bc6fcbf2b41e4eaebaf2410c500f21](https://github.com/AdaptiveCpp/AdaptiveCpp/commit/8a75ba2410bc6fcbf2b41e4eaebaf2410c500f21)
& HeCBench commit [8d66934f100d6d6c972ce476e7f67b3a0fdf0454](https://github.com/ORNL/HeCBench/commit/8d66934f100d6d6c972ce476e7f67b3a0fdf0454).*
# Under the Hood
Now that we’ve seen the backend in action, let’s find out how it works.
## Runtime Implementation
Before any ker