# The Case for Block-Based Programming With cuTile and Triton for Image Processing Workloads Used in Semi Manufacturing

## Abstract

Deploying image processing workloads on NVIDIA GPUs at near speed-of-light performance has required hand-tuning kernels using CUDA C++, given the unique nature of the loads used in semiconductor manufacturing. However, maintaining and debugging CUDA C++ code is challenging, often requiring significant re-tuning when adopting new GPU architectures. Tile-based programming, enabled by domain-specific languages (DSLs) such as NVIDIA cuTile and OpenAI Triton, offers a more portable and maintainable alternative while achieving performance close to hand-tuned CUDA C++. We'll evaluate tile-based programming for image processing workloads used in semiconductor manufacturing, comparing development effort and performance when porting operations from CUDA C++ to cuTile and Triton. We share insights from implementing GPU-accelerated algorithms, debugging, and optimizing performance in these DSLs.

## AI Summary

- The presentation covered the manufacturing process of semiconductors, emphasizing the complexity and the importance of process control to detect defects early.
- The company used a combination of inspection and metrology tools, including high-speed sensors and supercomputers, to ensure chip functionality.
- A comparison of CUDA, cuTile, and Triton for GPU programming was presented, focusing on performance, maintainability, and code simplicity.
- cuTile and Triton were found to offer significant reductions in lines of code compared to CUDA, making development and maintenance easier.
- For memory-bound operations, cuTile achieved high memory bandwidth, while Triton performed better in compute-bound scenarios.
- The programming experience with block-based languages was discussed, highlighting the benefits of higher-level abstractions and the challenges of limited low-level control.

## Transcript

**[00:00:10 – 00:00:13]** The talk is sort of split into three semi-equal parts,

**[00:00:13 – 00:00:14]** depending on how fast we go.

**[00:00:14 – 00:00:18]** We'll see which one takes more time, but I'll cover

**[00:00:18 – 00:00:21]** the first two parts where I'll really set the context up.

**[00:00:21 – 00:00:24]** I'll talk a little bit about how chips are manufactured.

**[00:00:24 – 00:00:26]** Now, this is a side of the world that most of us who

**[00:00:26 – 00:00:29]** design chips never think about, but it's a wonderful world

**[00:00:29 – 00:00:31]** of complexity, which...

**[00:00:31 – 00:00:33]** Requires something called process control to get right.

**[00:00:33 – 00:00:35]** So I'll talk a little bit about how chips are manufactured,

**[00:00:35 – 00:00:38]** what process control is, and then we'll get to the real challenge.

**[00:00:38 – 00:00:41]** Which is, we have this beast of a data processing stack

**[00:00:41 – 00:00:44]** and programming that is not a trivial task.

**[00:00:44 – 00:00:48]** And I'll then switch over to Arjun who will come and talk about how

**[00:00:48 – 00:00:51]** we've been exploring block-based programming with both Triton and

**[00:00:51 – 00:00:53]** Kutile on this particular stack.

**[00:00:54 – 00:00:58]** Again, as Rob mentioned, the both of us will be splitting this talk.

**[00:00:58 – 00:01:01]** We work in KLA's Artificial Intelligence and Advanced Computing

**[00:01:01 – 00:01:07]** Lab based out of the IIT Madras Research Park in Chennai, in India.

**[00:01:07 – 00:01:09]** Okay, so let's get started.

**[00:01:09 – 00:01:13]** The first part is to set the context, oh, it's weird to

**[00:01:13 – 00:01:14]** hear myself, but okay.

**[00:01:14 – 00:01:17]** The first part is to set the context, and I'll talk how

**[00:01:17 – 00:01:22]** semiconductors are manufactured and what actually process control is.

**[00:01:23 – 00:01:26]** So you see all around you these wonderful GPU devices

**[00:01:26 – 00:01:29]** that get designed, you see storage that gets designed, you see chips

**[00:01:29 – 00:01:34]** getting manufactured, but most of us who design chips only go to the

**[00:01:34 – 00:01:37]** point of designing the circuit and then you go into this black box and

**[00:01:37 – 00:01:39]** go manufacture it and come back.

**[00:01:39 – 00:01:41]** But if you really look under the hood and see what it takes to

**[00:01:41 – 00:01:46]** manufacture them, these chips today are built on top of transistors

**[00:01:46 – 00:01:51]** and NAND devices which are extremely complex in geometries.

**[00:01:51 – 00:01:55]** The picture on the left side that shows a schematic of what

**[00:01:55 – 00:01:57]** is a gate all around transistor.

**[00:01:57 – 00:01:59]** For those of you who study electrical engineering in college,

**[00:01:59 – 00:02:03]** we learned that a transistor is a source, a drain, and a gate.

**[00:02:03 – 00:02:05]** Life was very simple, but that was 50 years ago.

**[00:02:05 – 00:02:07]** This is what a transistor looks like today.

**[00:02:07 – 00:02:09]** If you ask me which one is the source, which one is the

**[00:02:09 – 00:02:12]** gate, and which one is the drain, it's literally gate is all around.

**[00:02:13 – 00:02:14]** That's what the name suggests.

**[00:02:14 – 00:02:18]** But what's important to realize is that the cross-section, the width

**[00:02:18 – 00:02:20]** of that gate or around transistor that I'm showing you here,

**[00:02:21 – 00:02:24]** is between 50 to 100 nanometers.

**[00:02:24 – 00:02:26]** And in physical dimensions, the cross-section of your

**[00:02:26 – 00:02:29]** hair is 10,000 nanometers.

**[00:02:29 – 00:02:32]** The cross-section of this beast is 50 to 100 nanometers.

**[00:02:32 – 00:02:34]** And look at the complexity of what you're trying to manufacture.

**[00:02:34 – 00:02:38]** The device in the middle is a schematic of what is 3D NAND

**[00:02:38 – 00:02:41]** today, which is the fundamental building block for NVMe, which

**[00:02:41 – 00:02:44]** is the storage in all your devices like your mobile phones

**[00:02:44 – 00:02:46]** and all your handhelds.

**[00:02:46 – 00:02:48]** Those aspect ratios that you're looking at is about

**[00:02:48 – 00:02:49]** 100 nanometers deep.

**[00:02:49 – 00:02:52]** You may say, hey, 100 nanometers is really tiny, but think

**[00:02:52 – 00:02:55]** about having to get that right.

**[00:02:55 – 00:02:57]** It needs to be perfectly parallel cylinders, and any mistake

**[00:02:57 – 00:03:00]** that you make, the entire chip will not work.

**[00:03:00 – 00:03:03]** The last picture that I have is one of what's called a reticle.

**[00:03:03 – 00:03:06]** A reticle is essentially what is the negative of the circuit

**[00:03:06 – 00:03:07]** you're trying to manufacture.

**[00:03:07 – 00:03:10]** And if you have a single mistake in the reticle, every chip

**[00:03:10 – 00:03:13]** that gets printed on every wafer that you manufacture

**[00:03:13 – 00:03:15]** for several months will be wrong.

**[00:03:15 – 00:03:18]** So getting that reticle perfectly right is extremely critical

**[00:03:18 – 00:03:21]** for getting the entire chip manufactured correctly.

**[00:03:21 – 00:03:24]** Now not only are these devices complex to manufacture, it

**[00:03:24 – 00:03:27]** also takes a long time.

**[00:03:27 – 00:03:30]** The typical chip to manufacture today takes a minimum of three

**[00:03:30 – 00:03:34]** months and the cartoon actually shows how a transistor grows,

**[00:03:34 – 00:03:37]** is cut, is sliced to actually get into the final NAND, the

**[00:03:38 – 00:03:41]** final gate-all-around transistor that I showed you earlier.

**[00:03:41 – 00:03:44]** And it takes anywhere between three to six months before which you

**[00:03:44 – 00:03:46]** can actually electrically test it.

**[00:03:46 – 00:03:48]** Only after that period, if you send a signal on one side,

**[00:03:48 – 00:03:51]** you can check, do I get a response on the other side?

**[00:03:51 – 00:03:55]** And if you made a mistake in that three to six months, let's say at

**[00:03:55 – 00:03:57]** the beginning of the first month,

**[00:03:57 – 00:03:59]** Everything that you manufacture through those months

**[00:03:59 – 00:04:00]** is not gonna work.

**[00:04:00 – 00:04:03]** So it's obviously extremely critical to make sure that

**[00:04:03 – 00:04:07]** the sooner you can check that the manufacturing is being done

**[00:04:07 – 00:04:10]** correctly, the higher chances that you'll get a functioning circuit.

**[00:04:10 – 00:04:12]** And that's really what KLA does.

**[00:04:12 – 00:04:15]** This process is called process control, where the goal is

**[00:04:15 – 00:04:17]** to make sure that...

**[00:04:17 – 00:04:21]** You find the defects before you actually electrically test

**[00:04:21 – 00:04:25]** them or you measure the transistors and the various dimensions of the

**[00:04:25 – 00:04:28]** transistors you're manufacturing before you find a fault in the

**[00:04:28 – 00:04:30]** device that's being manufactured.

**[00:04:30 – 00:04:32]** So, KLA makes machines or tools.

**[00:04:32 – 00:04:36]** A typical KLA tool is about half to a third of the size of the

**[00:04:36 – 00:04:38]** room that we're sitting in today.

**[00:04:38 – 00:04:41]** It's a big mass of machines that enable you to either do

**[00:04:41 – 00:04:43]** inspection, which is fault finding.

**[00:04:43 – 00:04:45]** You can see some of the pictures here.

**[00:04:45 – 00:04:47]** The first picture on the top left always blows my mind.

**[00:04:47 – 00:04:50]** It's a scanning electron microscope picture.

**[00:04:50 – 00:04:53]** The dimensions are about 40 by 40 nanometers.

**[00:04:53 – 00:04:55]** And you're really trying to find that little defect, right?

**[00:04:55 – 00:04:57]** It's really, really, really small to find.

**[00:04:57 – 00:04:59]** So we make machines that are inspectors that help you find

**[00:04:59 – 00:05:03]** these critical defects before they electrically can be tested

**[00:05:03 – 00:05:05]** or we make systems that are called metrology, shown on

**[00:05:05 – 00:05:08]** the right side, which enable you to measure critical dimensions.

**[00:05:09 – 00:05:11]** Again, going back to electrical engineering classes, critical

**[00:05:11 – 00:05:13]** dimension used to be distance between source and drain,

**[00:05:13 – 00:05:15]** very simple, but no longer.

**[00:05:16 – 00:05:19]** Typical transistors have something like 15 to 20 critical dimensions,

**[00:05:20 – 00:05:24]** including the angles, the depth of the well, the angles

**[00:05:24 – 00:05:27]** and the various dimensions that are important to get

**[00:05:27 – 00:05:29]** right for the entire thing to work.

**[00:05:29 – 00:05:31]** So that's what process control actually is.

**[00:05:31 – 00:05:34]** It's a means by which we do inspection and metrology to

**[00:05:34 – 00:05:38]** ensure that the chip finally actually works correctly.

**[00:05:38 – 00:05:42]** This is a picture of what a typical KLA tool looks like.

**[00:05:42 – 00:05:44]** It's a cyber-physical system, in my opinion, one of the

**[00:05:44 – 00:05:48]** most fascinating cyber-physical systems that has been built because

**[00:05:48 – 00:05:55]** it's really, it has everything to do with capturing the image,

**[00:05:55 – 00:05:57]** processing the image, the optics that generate the image, all of

**[00:05:58 – 00:06:00]** this is inside this physical tool.

**[00:06:00 – 00:06:03]** So, this particular tool has mechanics that robotic arms

**[00:06:03 – 00:06:06]** that actually go and place the wafer under the optics that's going

**[00:06:06 – 00:06:10]** to send the light, it has the light source that generates the light, it

**[00:06:10 – 00:06:13]** has a sensor that collects all the reflected and deflected light, and

**[00:06:14 – 00:06:17]** finally at the bottom, the image and data processing stack, which

**[00:06:17 – 00:06:20]** is really a supercomputer sitting physically inside this tool.

**[00:06:20 – 00:06:24]** So unlike many of the other industries that operate with

**[00:06:24 – 00:06:28]** cloud or data that's being processed elsewhere, we do

**[00:06:28 – 00:06:31]** edge computing in the sense that it's on device, except

**[00:06:31 – 00:06:33]** our device is a supercomputer.

**[00:06:33 – 00:06:35]** We're not talking about small cell phones, we're talking really

**[00:06:35 – 00:06:40]** about hundreds or thousands of GPUs sitting inside a physical tool.

**[00:06:40 – 00:06:44]** So the challenge really is how do you program this beast?

**[00:06:44 – 00:06:48]** The reason why it's important is because, again, we are shipping

**[00:06:48 – 00:06:51]** all the compute that you need to run these algorithms on the tool.

**[00:06:51 – 00:06:54]** So let's look at what that compute looks like.

**[00:06:54 – 00:06:56]** So this processing stack really, like I said, starts with optics.

**[00:06:56 – 00:06:59]** We have two different kinds of optics.

**[00:06:59 – 00:07:03]** Optics that are in the visible or in the optical range, which is

**[00:07:03 – 00:07:06]** 200 to 1,000 nanometers of light.

**[00:07:06 – 00:07:08]** We also have optics that are scanning electron microscope

**[00:07:08 – 00:07:12]** based, give you much better resolution, but are much slower.

**[00:07:12 – 00:07:14]** We then have these high-speed sensors that's collecting

**[00:07:14 – 00:07:17]** data at the rate of about 50 gigabytes per second.

**[00:07:17 – 00:07:21]** It's gigabytes per second, and so we generate petabytes

**[00:07:21 – 00:07:23]** of data every day, and we need to process through it.

**[00:07:23 – 00:07:25]** It's almost like an exoplanet search.

**[00:07:25 – 00:07:28]** Because you're really looking for that one small defect

**[00:07:28 – 00:07:31]** that's there somewhere, and then finally, the computing

**[00:07:31 – 00:07:34]** stack gets all this data, and what does that HPC stack look like?

**[00:07:34 – 00:07:37]** You get the raw input data that comes from the sensors.

**[00:07:37 – 00:07:39]** We then go through what I call optical detection.

**[00:07:39 – 00:07:42]** So the image there, or over there, it's not blurred.

**[00:07:42 – 00:07:45]** This is actually what an image from an optical inspector looks like.

**[00:07:46 – 00:07:48]** The resolution is very poor, because that's the only way

**[00:07:49 – 00:07:51]** for us to run it at speed.

**[00:07:51 – 00:07:54]** But the challenge really is, from this, we need to find

**[00:07:54 – 00:07:58]** out if that blink is actually an unresolved mistake, that

**[00:07:58 – 00:08:00]** is, there's no defect, or is there actually a defect?

**[00:08:00 – 00:08:04]** So what our algorithms do is, from that optical image, it tries

**[00:08:04 – 00:08:08]** to decipher and predict, is there a defect or is there no defect?

**[00:08:09 – 00:08:11]** And the process engineer actually goes to the scanning electron

**[00:08:11 – 00:08:14]** microscope for verification.

**[00:08:14 – 00:08:16]** He or she goes and puts it under the scanning electron microscope.

**[00:08:16 – 00:08:18]** That gives you a beautifully resolved image.

**[00:08:18 – 00:08:20]** It's actually the same image.

**[00:08:20 – 00:08:21]** You can see what the optical image looks like.

**[00:08:21 – 00:08:23]** I don't think with our naked eye we can find

**[00:08:23 – 00:08:25]** that there's a defect or not.

**[00:08:25 – 00:08:29]** But the goal of our algorithms is to predict that, is that

**[00:08:29 – 00:08:31]** when it predicts that there's a defect, the SEM actually says,

**[00:08:31 – 00:08:34]** yes, there indeed is a defect.

**[00:08:34 – 00:08:37]** And finally, what we do is we classify the faults and we give

**[00:08:37 – 00:08:42]** it out as a Pareto for the process engineer to look at what classes of

**[00:08:42 – 00:08:45]** faults are being found so that they can go fix their process tools to

**[00:08:45 – 00:08:46]** manufacture the devices correctly.

**[00:08:46 – 00:08:50]** And this goes on and on until you get sufficient yield.

**[00:08:50 – 00:08:53]** In terms of data processing, we have a combination of

**[00:08:53 – 00:08:56]** both real-time and bulk processing happening within

**[00:08:56 – 00:08:58]** the same physical device.

**[00:08:58 – 00:09:01]** So we are neither throughput nor latency-prone, we are

**[00:09:01 – 00:09:03]** both, depending on what part of the algorithm, depending

**[00:09:03 – 00:09:07]** on what part of the HPC stack you're actually working with.

**[00:09:07 – 00:09:10]** So we've been collaborating very closely with NVIDIA for

**[00:09:10 – 00:09:14]** close to eight years now to try and move a lot of these workloads that

**[00:09:14 – 00:09:16]** used to run on the CPU to the GPU.

**[00:09:16 – 00:09:17]** It's been a wonderful collaboration.

**[00:09:17 – 00:09:19]** We've had,

**[00:09:20 – 00:09:22]** Loads that are more matrix-matrix like, shown on top, which

**[00:09:22 – 00:09:25]** are typical inspection workloads where we've gotten very good gains.

**[00:09:25 – 00:09:30]** But there are also workloads today that are vector workloads.

**[00:09:30 – 00:09:33]** Most of our metrology workloads tend to work on more sparse

**[00:09:33 – 00:09:36]** data, and therefore they tend to be more vector-like, and

**[00:09:36 – 00:09:39]** as much as possible, of course, we have tried to imbibe various types

**[00:09:39 – 00:09:43]** of AI inside this entire process.

**[00:09:43 – 00:09:46]** Now, throughout this, one of the core driving forces

**[00:09:46 – 00:09:51]** for us to decide how to migrate these algorithms to the GPU

**[00:09:51 – 00:09:54]** has been guided by what we call the speed of light model.

**[00:09:54 – 00:09:55]** So let me just quickly walk you through this.

**[00:09:55 – 00:09:58]** It's a simple analytical model considering a workload that's

**[00:09:58 – 00:10:02]** comprised of N compute operations and M memory operations.

**[00:10:02 – 00:10:05]** We do a simple analytical model where we say, we find

**[00:10:05 – 00:10:08]** the amount of time we expect the computation to take and

**[00:10:08 – 00:10:11]** the memory operations to take just by dividing it with a throughput.

**[00:10:11 – 00:10:14]** Then depending on whether you're a latency use case or a throughput

**[00:10:14 – 00:10:19]** use case, the best speed of light model will either be a sum of TComp

**[00:10:19 – 00:10:24]** and Tmem or will be a max of TComp and Tmem if you're in throughput.

**[00:10:24 – 00:10:28]** Now this is an analytical model and this is not the

**[00:10:28 – 00:10:29]** roofline plot that you see.

**[00:10:29 – 00:10:32]** And we find this analytical model to be a better guiding source

**[00:10:32 – 00:10:37]** for our optimizations because it is not colored by the implementation.

**[00:10:37 – 00:10:40]** In your roofline model, if you started with a bad implementation,

**[00:10:40 – 00:10:42]** you get automatically colored by it, but this is actually

**[00:10:42 – 00:10:45]** saying truly speed of light, assuming there are no more

**[00:10:45 – 00:10:47]** bottlenecks, which is, of course, not the way the real world is,

**[00:10:47 – 00:10:50]** what you actually expect to see.

**[00:10:50 – 00:10:53]** The key question that we ask, of course, is, does your algorithm

**[00:10:53 – 00:10:56]** achieve speed of light on the entitled hardware for

**[00:10:56 – 00:10:57]** the given algorithm?

**[00:10:57 – 00:11:00]** And we find the right abstraction layer such that we can actually

**[00:11:00 – 00:11:03]** achieve the speed of light.

**[00:11:03 – 00:11:06]** So we started out with this journey by looking very close

**[00:11:06 – 00:11:08]** to metal and we said, okay, let's look at a few different

**[00:11:08 – 00:11:13]** GPUs and look at their GPU-native languages, CUDA for NVIDIA, HIP for

**[00:11:13 – 00:11:19]** AMD, and DPC++ for Intel, and we said, let's find out which hardware

**[00:11:19 – 00:11:23]** has the best GPU-native software such that we can actually program

**[00:11:23 – 00:11:25]** with it to get speed of light.

**[00:11:25 – 00:11:28]** Remember, it's not fair to compare absolute runtimes

**[00:11:28 – 00:11:30]** for different workloads on different hardware, because

**[00:11:30 – 00:11:33]** the speeds and feeds your computer and memory could be different.

**[00:11:33 – 00:11:35]** So your absolute runtime might not just compare directly.

**[00:11:35 – 00:11:38]** A speed-of-flight model predicts the absolute runtime, that's

**[00:11:38 – 00:11:39]** a better comparison.

**[00:11:39 – 00:11:42]** So there's some work we published a while back which actually showed

**[00:11:42 – 00:11:47]** that on different, on what we found here was that on NVIDIA hardware

**[00:11:47 – 00:11:51]** run here on the A10 GPU, The CPU, the CUDA runtime was actually

**[00:11:51 – 00:11:55]** much closer to the predicted.

**[00:11:55 – 00:11:57]** Throughput speed of light with throughput computation

**[00:11:57 – 00:11:58]** than the other hardware.

**[00:11:58 – 00:12:02]** This resulted in us investing a lot in NVIDIA hardware and

**[00:12:02 – 00:12:06]** looking at optimizing with low-level CUDA pretty extensively.

**[00:12:06 – 00:12:07]** Now, this has been fun, right?

**[00:12:07 – 00:12:09]** Those of us who are sitting here who've written CUDA programs, it's

**[00:12:09 – 00:12:14]** a lot of fun to write CUDA, but a three-line Torch code translates

**[00:12:14 – 00:12:17]** to a 185-line CUDA kernel.

**[00:12:18 – 00:12:20]** And the unique challenge for us in this, is that not only

**[00:12:20 – 00:12:24]** does it take time to develop, but our tools go into the fab and

**[00:12:24 – 00:12:28]** live in the fab for 15, 20 years.

**[00:12:28 – 00:12:31]** And if I were to go back and debug that line of code,

**[00:12:31 – 00:12:34]** Two years down the line, 10 years down the line, I don't

**[00:12:34 – 00:12:35]** think I want to do that.

**[00:12:35 – 00:12:37]** So this is becoming increasingly difficult

**[00:12:37 – 00:12:38]** for us to get performance.

**[00:12:38 – 00:12:41]** So we've been trying to work on abstractions that give

**[00:12:41 – 00:12:45]** us better programmability while getting the same performance, and

**[00:12:45 – 00:12:48]** block-based programming is really, the fact that it's gaining momentum

**[00:12:48 – 00:12:51]** has caught our attention a lot.

**[00:12:51 – 00:12:53]** So this is a slide, thanks to Steven, who had introduced

**[00:12:53 – 00:12:57]** this at GTC last year, which actually shows the fundamental idea

**[00:12:57 – 00:12:59]** behind block-based programming.

**[00:12:59 – 00:13:01]** When you write in PyTorch, which is grid-based programming, you just

**[00:13:01 – 00:13:04]** look at an image as a whole image and you just operate on the image.

**[00:13:04 – 00:13:06]** On the right extreme, you have thread-level programming,

**[00:13:06 – 00:13:09]** where you do CUDA, CUDA C++, where you operate on threads

**[00:13:09 – 00:13:11]** and you look at literally every pixel and say, what does the

**[00:13:11 – 00:13:14]** thread that operates on a pixel do?

**[00:13:14 – 00:13:17]** The midpoint, which is block-based programming, enables you as

**[00:13:17 – 00:13:21]** a programmer to break the grid into a bunch of blocks

**[00:13:21 – 00:13:24]** and lets the compiler actually translate from the block

**[00:13:24 – 00:13:29]** And therefore, it makes it easier for the compiler to

**[00:13:29 – 00:13:31]** actually arguably extract performance and makes application

**[00:13:31 – 00:13:35]** writer work also easy.

**[00:13:35 – 00:13:38]** On the right side, the graph is a schematic shown by Phil Tilley

**[00:13:38 – 00:13:41]** in the Triton Conference last year, where we've seen that over

**[00:13:41 – 00:13:45]** the last six months to 12 months, there's been a burst in the number

**[00:13:46 – 00:13:49]** of different DSLs and Python-based DSLs that really enable you

**[00:13:49 – 00:13:53]** to walk on this curve between performance and productivity.

**[00:13:53 – 00:13:55]** On the top left is...

**[00:13:55 – 00:13:58]** The SAS write, for those of us who dare to write all the way down

**[00:13:58 – 00:14:01]** in SAS, a little more model-ish, CUDA is right next to that.

**[00:14:01 – 00:14:04]** And the bottom right side is writing fully in numpy.

**[00:14:04 – 00:14:06]** And there's a lot of stops along that path.

**[00:14:06 – 00:14:08]** And depending on where you want to land, you can actually

**[00:14:08 – 00:14:09]** go and play with it.

**[00:14:09 – 00:14:11]** And the one thing that really interests us is the fact

**[00:14:11 – 00:14:15]** that, for images, block-based programming is very natural.

**[00:14:15 – 00:14:17]** Because you always think about images in blocks.

**[00:14:17 – 00:14:19]** And depending on which block you select, you can see that,

**[00:14:19 – 00:14:23]** in this image, the semantics of the image is also changing.

**[00:14:23 – 00:14:25]** So you can actually do a lot of control with it.

**[00:14:25 – 00:14:27]** So we came back to the same question and we said, okay,

**[00:14:27 – 00:14:30]** block-based is interesting, let's start looking at that, but

**[00:14:30 – 00:14:33]** can we achieve the same speed of light for image processing kernels

**[00:14:33 – 00:14:35]** on GPUs as we did with CUDA?

**[00:14:36 – 00:14:40]** And are these kernels actually easy to develop and to maintain

**[00:14:41 – 00:14:44]** while compared to the GPU native languages?

**[00:14:44 – 00:14:47]** So to talk more about that, let me switch over to Arjun, who'll

**[00:14:47 – 00:14:50]** talk about our evaluation of this.

**[00:14:50 – 00:14:53]** So what we're going to do now is we look at some of

**[00:14:53 – 00:14:56]** the workloads that we run at KLA.

**[00:14:56 – 00:15:00]** I'll walk you through how we might express that in block-based

**[00:15:00 – 00:15:04]** programming languages, the performance we get from that,

**[00:15:04 – 00:15:08]** some optimizations that you can fold in, and then I'll sum

**[00:15:08 – 00:15:15]** it up with the takeaways and what the programming experience is like.

**[00:15:15 – 00:15:19]** So first, a quick primer on QTile and Triton.

**[00:15:19 – 00:15:22]** So this is a slide that we borrowed from Steven's talk last

**[00:15:22 – 00:15:25]** year, where he introduced QTile.

**[00:15:25 – 00:15:30]** On the left is a softmax function in numpy, and on the right

**[00:15:30 – 00:15:34]** is a softmax kernel in QTile.

**[00:15:34 – 00:15:39]** The key difference here is that now you have a Tile load and

**[00:15:39 – 00:15:44]** a Tile store, and once you load it in, it's pretty much the same set

**[00:15:44 – 00:15:48]** of operations that you would do in numpy, and then you store it back.

**[00:15:48 – 00:15:55]** So the Python front end for QTile is built on top of this

**[00:15:55 – 00:15:58]** Tile IR extraction.

**[00:15:58 – 00:16:02]** So TILE-IR is, you can think of it as a virtual ISA for

**[00:16:02 – 00:16:07]** TILE-based programming languages, analogous to how PTX is a

**[00:16:07 – 00:16:10]** virtual ISA for SIMT programs.

**[00:16:10 – 00:16:13]** And the mapping from TILE-IR down to SAS is handled by

**[00:16:13 – 00:16:18]** a proprietary MLIR-based compiler called TILE-IR-AS.

**[00:16:19 – 00:16:24]** Bryce Lebeck gave a wonderful talk just the last hour, if you missed

**[00:16:24 – 00:16:28]** it I recommend revisiting it, where he got into some of the details of

**[00:16:28 – 00:16:32]** how you might write QTile kernels.

**[00:16:33 – 00:16:36]** What's interesting for us is also the interoperability between

**[00:16:36 – 00:16:41]** QTile and CUDA C++ functions, so this means that you don't have to

**[00:16:42 – 00:16:45]** Invest fully into either the tile-based or the thread-based

**[00:16:45 – 00:16:48]** programming paradigms, you can actually pick and choose

**[00:16:48 – 00:16:51]** for the operations as you wish.

**[00:16:52 – 00:16:55]** Looking at Triton, Triton is really what brought

**[00:16:55 – 00:16:59]** block-based GPU programming to the mainstream, and today

**[00:16:59 – 00:17:03]** it sits as a first-class citizen in PyTorch's inductor-compiler flow.

**[00:17:03 – 00:17:08]** A lot of your gem kernels are fused activations and attention

**[00:17:08 – 00:17:11]** kernels that you would get if you were to express them with

**[00:17:11 – 00:17:18]** PyTorch, flows through Triton and picks templates written in Triton.

**[00:17:19 – 00:17:25]** Triton also is built on top of an MLIR compiler with public back-ins

**[00:17:25 – 00:17:29]** for both NVIDIA and AMD GPUs.

**[00:17:29 – 00:17:32]** This figure here highlights the contrast between the thread-based

**[00:17:32 – 00:17:36]** programming paradigm and Block-based programming paradigm.

**[00:17:36 – 00:17:39]** While in thread-based programming, you actually express what

**[00:17:39 – 00:17:44]** each thread does, the operations that each thread does, and

**[00:17:44 – 00:17:48]** the data that it owns, when it comes to block-based programming,

**[00:17:48 – 00:17:52]** you just express your algorithm as a set of operations on

**[00:17:52 – 00:17:57]** tiles, and you rely on a compiler to map this down to the threads.

**[00:17:57 – 00:17:59]** Triton's compiler stack goes through multiple

**[00:17:59 – 00:18:01]** levels of abstraction.

**[00:18:01 – 00:18:06]** The first, Triton IR, or TTIR, is hardware agnostic.

**[00:18:06 – 00:18:09]** This is where you've got some hardware agnostic

**[00:18:09 – 00:18:11]** abstractions and passes.

**[00:18:11 – 00:18:15]** Triton GPIO IR implements some NVIDIA-specific or

**[00:18:15 – 00:18:17]** AMD-specific passes.

**[00:18:17 – 00:18:20]** And this is where some of the abstractions that someone

**[00:18:20 – 00:18:23]** familiar with CUDA would...

**[00:18:24 – 00:18:27]** You then go to LLVM and from there you can branch out to

**[00:18:27 – 00:18:30]** either PTX or AMD's source code.

**[00:18:30 – 00:18:35]** Triton also now supports compiling from Triton IR down to Tile IR

**[00:18:35 – 00:18:37]** via the Triton to Tile IR backend.

**[00:18:37 – 00:18:40]** This is a public open source project in Triton's

**[00:18:40 – 00:18:42]** GitHub repository.

**[00:18:42 – 00:18:44]** And once you've generated TILE-IR, you can actually go

**[00:18:44 – 00:18:50]** down to Kube in SAS via TILE-IR AS.

**[00:18:50 – 00:18:51]** Okay, so now for our evaluation.

**[00:18:51 – 00:18:55]** We actually revisit the questions that Pradeep posed.

**[00:18:55 – 00:18:58]** The questions are, can you get performance near speed

**[00:18:58 – 00:19:02]** of light for image processing workloads when you express it in

**[00:19:02 – 00:19:04]** block-based programming languages?

**[00:19:04 – 00:19:07]** And two, in doing so, can you still write code that

**[00:19:07 – 00:19:09]** is maintainable and modular?

**[00:19:10 – 00:19:12]** We base this evaluation on

**[00:19:12 – 00:19:14]** Operations from three classes.

**[00:19:14 – 00:19:17]** First is point-wise and reduction operations.

**[00:19:17 – 00:19:21]** These are inherently data parallel operations with regular

**[00:19:21 – 00:19:24]** data access patterns and regular compute patterns.

**[00:19:24 – 00:19:28]** The compute operations are typically floating-point arithmetic

**[00:19:28 – 00:19:31]** running on CUDA cores, and they tend to be DRAM-bound

**[00:19:31 – 00:19:33]** operations in general.

**[00:19:33 – 00:19:36]** The second class of operations are gather-scatter So these

**[00:19:36 – 00:19:37]** are some of the major operations.

**[00:19:37 – 00:19:40]** Where you might have irregular data access patterns, one such

**[00:19:40 – 00:19:45]** example being pointer-chasing data accesses like what I've shown here,

**[00:19:45 – 00:19:48]** or you could have irregular compute patterns where you've got some

**[00:19:48 – 00:19:52]** atomic contention or some imbalance in the work that you do for thread.

**[00:19:53 – 00:19:56]** The third class that we look at are tensor operations,

**[00:19:56 – 00:19:58]** and in particular, we look at 2D convolutions.

**[00:19:59 – 00:20:03]** Why convolutions are interesting is because these can be mapped

**[00:20:03 – 00:20:05]** to gem operations on Tensor Cores.

**[00:20:05 – 00:20:08]** But to do so, you need to do some data marshaling, in particular,

**[00:20:08 – 00:20:11]** the IMT call transformation.

**[00:20:11 – 00:20:16]** On the input, now how cleverly you fold this data marshaling

**[00:20:16 – 00:20:20]** in dictates whether or not your kernel is going to be

**[00:20:20 – 00:20:24]** compute bound and how effectively it uses a Tensor Cores throughput.

**[00:20:24 – 00:20:27]** If you incur latencies in the data marshaling, you actually

**[00:20:27 – 00:20:30]** take a toll on the throughput.

**[00:20:30 – 00:20:34]** So we wanted to see how effective these languages are when it comes

**[00:20:34 – 00:20:37]** to expressing these operations.

**[00:20:39 – 00:20:43]** Okay, so for the first class of operations, these are inherently

**[00:20:43 – 00:20:46]** data parallel operations.

**[00:20:46 – 00:20:49]** Some examples of this would be Gaussian filters, Stencil

**[00:20:49 – 00:20:55]** computations, global reduction on an image, mass sum, or

**[00:20:55 – 00:20:56]** 2D interpolation.

**[00:20:56 – 00:20:59]** Now because these are inherently data parallel,

**[00:20:59 – 00:21:02]** you can actually chunk the data that sits and Yadira.

**[00:21:02 – 00:21:05]** Load in a tile into your thread block scope.

**[00:21:05 – 00:21:09]** You might have multiple inputs, so you can have multiple input tiles.

**[00:21:09 – 00:21:12]** Perform the floating-point arithmetic on them, generate

**[00:21:12 – 00:21:15]** the output tile, and then write it back to DRAM.

**[00:21:16 – 00:21:18]** A CUDA programmer gets to see the entire memory hierarchy

**[00:21:18 – 00:21:22]** of the GPU, so that's your caches, your shared memory, registers,

**[00:21:22 – 00:21:26]** and if you're on a big blackwell, you also get to see tensor memory.

**[00:21:26 – 00:21:29]** However, in block-based programming languages, you only get to

**[00:21:29 – 00:21:32]** see two scopes, the DRAM scope and the tile scope.

**[00:21:32 – 00:21:35]** Now, the data that resides in the tile scope, whether

**[00:21:35 – 00:21:40]** or not it maps, which element in the memory hierarchy it maps

**[00:21:40 – 00:21:43]** to, is decided by the compiler.

**[00:21:43 – 00:21:47]** What we want is for these tiles to be staged in shared memory,

**[00:21:47 – 00:21:52]** so that you can actually issue asynchronous loads using TMA or the

**[00:21:52 – 00:21:55]** CP.async constructions, And then...

**[00:21:55 – 00:21:59]** Schedule them using software pipeline schedules so that

**[00:21:59 – 00:22:04]** you have compute and data access happening concurrently.

**[00:22:04 – 00:22:08]** Now, because these are memory-bound operations, we use a speed

**[00:22:08 – 00:22:12]** of light estimate, which is given by the total memory traffic

**[00:22:12 – 00:22:16]** divided by the DRAM bandwidth.

**[00:22:17 – 00:22:23]** So these are results that we got on an RTX 1590 with CUDA 13.1.

**[00:22:23 – 00:22:24]** There are three bars here.

**[00:22:24 – 00:22:29]** The first, the green bar, is for kernels written in Qtile Python.

**[00:22:29 – 00:22:32]** The black bar are for kernels written in Triton and compiled

**[00:22:32 – 00:22:35]** via Triton's stock compiler flow.

**[00:22:36 – 00:22:38]** And the gray bars are the same Triton kernels when compiled

**[00:22:38 – 00:22:41]** via Triton's TileLayer backend.

**[00:22:42 – 00:22:44]** On the x-axis, you've got different operations, and on

**[00:22:44 – 00:22:47]** the y-axis, you've got performance relative to speed of light.

**[00:22:47 – 00:22:50]** Higher is better, and you max out at 1.

**[00:22:50 – 00:22:54]** If you go beyond 1, you're faster than light, you break relativity,

**[00:22:54 – 00:22:57]** and then physicists panic, right?

**[00:22:58 – 00:23:02]** So I want to highlight two key results from this chart.

**[00:23:02 – 00:23:06]** The first is that for Gaussian filters, Stencil at small

**[00:23:06 – 00:23:10]** radius, and interpolation, you actually achieve a considerably

**[00:23:10 – 00:23:15]** high memory bandwidth, upwards of 80% with QTile.

**[00:23:15 – 00:23:18]** This is really interesting, and in fact we noticed that for

**[00:23:18 – 00:23:25]** some of these, we were even able to use the TMA for loads and stores.

**[00:23:25 – 00:23:27]** So that's quite exciting.

**[00:23:27 – 00:23:29]** The second trend that I want to highlight is of course

**[00:23:29 – 00:23:31]** the elephant in the room.

**[00:23:31 – 00:23:35]** At higher ADI, for the stencil operations, your

**[00:23:35 – 00:23:37]** performance takes a toll.

**[00:23:37 – 00:23:40]** This is true for both QTile and Triton.

**[00:23:40 – 00:23:42]** However, it's more pronounced for Triton.

**[00:23:42 – 00:23:45]** And one thing that I want to highlight here, so the reason

**[00:23:45 – 00:23:50]** why there's a drop in performance is because one would want explicit

**[00:23:50 – 00:23:55]** control over the data residing in shared memory, so that regardless

**[00:23:55 – 00:23:57]** of the radius you have...

**[00:23:57 – 00:24:00]** Fast access to the data, and you can compute the reduction

**[00:24:00 – 00:24:05]** in your CUDA course, but since you don't have that level of

**[00:24:05 – 00:24:10]** control in block-based programming languages, you do incur some cache

**[00:24:10 – 00:24:12]** Cashness is at the end of the day.

**[00:24:12 – 00:24:17]** Now interestingly, with TILE-IR, the cache misses that you incur is

**[00:24:17 – 00:24:19]** less pronounced than with Triton.

**[00:24:19 – 00:24:24]** So regardless of the fact that you wrote the Triton to TILE-IR

**[00:24:24 – 00:24:28]** bar in Triton, because you're using the TILE-IR backend, it's actually

**[00:24:28 – 00:24:32]** able to manage the cache hits better, and you get performance

**[00:24:32 – 00:24:38]** on par with QTile Python, and about 2x that of Triton.

**[00:24:39 – 00:24:43]** We also look at the number of lines of code as a proxy for

**[00:24:43 – 00:24:46]** how simple or complex a kernel is.

**[00:24:46 – 00:24:50]** The general trend that you would find is that you get,

**[00:24:50 – 00:24:54]** you can express the operations at about 50% of the lines

**[00:24:54 – 00:24:56]** of code of a CUDA kernel.

**[00:24:56 – 00:24:59]** So for instance, a 300 line Gaussian filter translates

**[00:24:59 – 00:25:04]** to about 100 to 150 lines of code in CUDA or Triton.

**[00:25:04 – 00:25:07]** You might notice an anomaly for the last row, where the

**[00:25:07 – 00:25:12]** Triton implementation for 2D interpolation is higher, uses more

**[00:25:12 – 00:25:14]** lines of code than a CUDA kernel.

**[00:25:14 – 00:25:18]** This actually has to do with the fact that in Triton, you

**[00:25:18 – 00:25:22]** pass in the offsets, and you have a pointer-based load semantic.

**[00:25:23 – 00:25:27]** This leads to a slightly verbose code for Triton operations,

**[00:25:27 – 00:25:32]** and it can actually cause There are some code blocks.

**[00:25:33 – 00:25:38]** The second class of operations that we look at are gather-scatter

**[00:25:38 – 00:25:42]** operations, and we use a case study with k-means clustering.

**[00:25:42 – 00:25:44]** This involves multiple stages.

**[00:25:44 – 00:25:48]** What's relevant here are the feature-prone centroid

**[00:25:48 – 00:25:51]** sample and k-means update stages, where you've got

**[00:25:51 – 00:25:54]** some gather-scatter-like accesses.

**[00:25:54 – 00:25:57]** So in case of feature-prone

**[00:25:57 – 00:26:02]** You have scattered accesses where you pick only some

**[00:26:02 – 00:26:06]** relevant data elements from a feature map and populate

**[00:26:06 – 00:26:08]** it into the proven features.

**[00:26:08 – 00:26:11]** In case of centroid sample, you have some random number

**[00:26:11 – 00:26:16]** generation, some sorting on a score function, and some atomic updates.

**[00:26:16 – 00:26:19]** And similarly for K-Means update, you have some

**[00:26:19 – 00:26:22]** Distance computation and some atomic updates as well.

**[00:26:23 – 00:26:28]** Now if you had control over the kernels, if you had complete

**[00:26:28 – 00:26:30]** control over the kernels, these are the optimizations

**[00:26:30 – 00:26:32]** that we would want to use.

**[00:26:32 – 00:26:34]** First would be kernel fusions so that you reduce the number

**[00:26:34 – 00:26:36]** of round trips to memory.

**[00:26:36 – 00:26:39]** Second would be to perform the atomic updates to shared

**[00:26:39 – 00:26:43]** memory so that you reduce the contention at L2 cache.

**[00:26:43 – 00:26:46]** And third, because you've got this catalytic axis, you inevitably

**[00:26:46 – 00:26:49]** have to use ct.where operations.

**[00:26:49 – 00:26:54]** So you try to reduce or eliminate them wherever possible.

**[00:26:56 – 00:27:00]** These, again, we look at QTile, Triton, and Triton

**[00:27:00 – 00:27:01]** compiled with TileIR.

**[00:27:01 – 00:27:04]** However, unlike the last time, we are not looking at performance

**[00:27:04 – 00:27:06]** relative to speed of light.

**[00:27:06 – 00:27:09]** We are looking at performance relative to Triton.

**[00:27:09 – 00:27:11]** The reason for this is because we used slightly different

**[00:27:11 – 00:27:15]** decompositions for QTile and Triton, and because this is

**[00:27:15 – 00:27:19]** inherently an iterative algorithm, coming up with a standard speed

**[00:27:19 – 00:27:20]** of light was not straightforward.

**[00:27:20 – 00:27:23]** So if you notice, the Triton bars are all at 1.

**[00:27:24 – 00:27:29]** And the bars for QTile and TritonTile are relative to this.

**[00:27:29 – 00:27:31]** Again, I want to highlight two key trends.

**[00:27:31 – 00:27:35]** First is that for k-means update, we found the bottleneck to be

**[00:27:35 – 00:27:39]** in the distance computation kernel, and we were able to actually

**[00:27:39 – 00:27:43]** write a better implementation in QTile, which gives us substantially

**[00:27:43 – 00:27:46]** better performance.

**[00:27:47 – 00:27:50]** The second trend is for centroid sampling, there's a random

**[00:27:50 – 00:27:53]** number generation step involved.

**[00:27:53 – 00:27:56]** QTile today does not support random number generation, so

**[00:27:56 – 00:28:00]** we actually had to use torch.rand, and that incurs some latency.

**[00:28:00 – 00:28:04]** In addition to that, you actually cannot use atomics,

**[00:28:04 – 00:28:06]** you cannot perform atomic updates to shared memory

**[00:28:06 – 00:28:08]** in Tile IR as it stands today.

**[00:28:08 – 00:28:11]** So you perform these atomic updates to global memory,

**[00:28:11 – 00:28:16]** which again causes higher contention End slows down a bit.

**[00:28:17 – 00:28:20]** In terms of number of lines of code, the story is similar.

**[00:28:20 – 00:28:27]** You're about at 50% of the lines of code of CUDA with QTile and Triton.

**[00:28:27 – 00:28:30]** Now, the third would be something that is arguably the most

**[00:28:31 – 00:28:35]** exciting class of operations here, tensicle convolutions.

**[00:28:35 – 00:28:41]** So to quickly recap, you've got an image, HWC, you've got C

**[00:28:41 – 00:28:45]** channels of an input, and you want to convolve it with K filters, each

**[00:28:45 – 00:28:50]** filter of shape R cross S, and one filter Interprets input channel.

**[00:28:50 – 00:28:55]** What you want to do is to load the filters to your TileScope

**[00:28:55 – 00:29:00]** and stream the inputs as blocks to TileScope, perform IM2

**[00:29:00 – 00:29:04]** call transformation implicitly, and convert it into a matrix-matrix

**[00:29:04 – 00:29:09]** product, feed the Tensor Cores, accumulate the results from

**[00:29:09 – 00:29:12]** Tensor Cores, and write the data back to DRAM.

**[00:29:12 – 00:29:15]** What are some optimizations that you would want to use here?

**[00:29:15 – 00:29:19]** First, to stream the inputs, you leverage asynchronous loads.

**[00:29:19 – 00:29:22]** Through the TMA or CP.async.

**[00:29:22 – 00:29:27]** For implicit IAM to call, you use shifted loads, so

**[00:29:27 – 00:29:30]** you work out a way to extract a sub-tile from a tile.

**[00:29:31 – 00:29:33]** Interestingly, QTile does have an API that supports

**[00:29:33 – 00:29:36]** this, called ct.extract.

**[00:29:37 – 00:29:41]** And finally, you use an asynchronous store, so that

**[00:29:41 – 00:29:46]** you can now have a three-stage pipeline to do your kernel.

**[00:29:46 – 00:29:51]** We ran into a few limitations with this operation here.

**[00:29:51 – 00:29:55]** First is that the ct.extractOp, while it is supported, is

**[00:29:55 – 00:29:58]** limited to subtitling only at boundaries of powers of

**[00:29:58 – 00:30:02]** two, which means if you want to implement a stride one convolution,

**[00:30:03 – 00:30:10]** you shift by one, an offset of one after each iteration, and since

**[00:30:11 – 00:30:13]** that wouldn't lead to boundaries that are powers of two, you

**[00:30:13 – 00:30:17]** cannot use the extract operation.

**[00:30:17 – 00:30:21]** As a result, we had to actually fall back to doing the IAM2CALL

**[00:30:21 – 00:30:24]** transformation from global memory, which also meant that

**[00:30:24 – 00:30:27]** we were not able to use the TMA.

**[00:30:29 – 00:30:32]** For this, again, we look at QTile, Triton, and TritonTileAR.

**[00:30:32 – 00:30:35]** The performance is relative to speed of light.

**[00:30:35 – 00:30:38]** And on the x-axis, you've got different channel configurations.

**[00:30:38 – 00:30:41]** So you've got different input channels and

**[00:30:41 – 00:30:42]** correspondingly output channels.

**[00:30:42 – 00:30:45]** The first three configurations are memory-bound on an RTX

**[00:30:45 – 00:30:49]** 5090, while the remaining four are compute-bound.

**[00:30:49 – 00:30:52]** What we find is that for the memory-bound kernels, you

**[00:30:52 – 00:30:56]** have one implementation that at least hits 60% of your

**[00:30:56 – 00:30:59]** bandwidth, which is quite good.

**[00:30:59 – 00:31:02]** Whereas when you go to compute-bound regime,

**[00:31:02 – 00:31:07]** you sort of max out at 50% of the compute flops.

**[00:31:07 – 00:31:10]** Interestingly, Triton actually does better than both QTile

**[00:31:10 – 00:31:14]** and Triton TileAR consistently for all of these operations.

**[00:31:14 – 00:31:16]** And we were actually able to use the TMA in some of

**[00:31:16 – 00:31:19]** the cases with Triton.

**[00:31:20 – 00:31:23]** In terms of lines of code, the CUDA kernel that one would

**[00:31:23 – 00:31:27]** write for this is at about 500 lines, and it's got CUDA and

**[00:31:27 – 00:31:32]** PTX, and the QTIL or Triton kernel is at about 80 lines of code.

**[00:31:32 – 00:31:36]** Now, I should admit, if you've got a GPU, the 500 lines of

**[00:31:36 – 00:31:39]** CUDA and PTX code that you would write for a convolution

**[00:31:39 – 00:31:43]** is one of the most exciting things you can do, but debugging those 500

**[00:31:43 – 00:31:44]** lines of code, not so much, right?

**[00:31:44 – 00:31:49]** So 80 lines of code, it's a really big boom.

**[00:31:49 – 00:31:53]** Okay, so now I'm going to switch gears and talk about

**[00:31:53 – 00:31:56]** the programming experience.

**[00:31:56 – 00:31:58]** First I'm going to talk about a comparison of CUDA Tile and

**[00:31:58 – 00:32:01]** Triton, looking at three aspects.

**[00:32:01 – 00:32:04]** First is the Python front end, second is the middle

**[00:32:04 – 00:32:08]** end, or the compiler stack, Tile IR or Triton IR, and

**[00:32:08 – 00:32:11]** third is performance debugging.

**[00:32:11 – 00:32:15]** When it comes to the Python frontend, Qtile actually supports

**[00:32:15 – 00:32:19]** automatic mapping of your loads and stores to TMA.

**[00:32:19 – 00:32:22]** You don't have to worry about tensor descriptors and details

**[00:32:22 – 00:32:27]** related to that, which is an interesting thing.

**[00:32:28 – 00:32:31]** We find that as of CUDA 13.1, Triton supports a

**[00:32:31 – 00:32:35]** richer set of operations, although some of those gaps

**[00:32:35 – 00:32:38]** have been bridged with CUDA 13.2.

**[00:32:38 – 00:32:40]** The third point is also quite interesting.

**[00:32:41 – 00:32:44]** Qtile has a distinction between a kernel and a device

**[00:32:44 – 00:32:47]** function analogous to the same distinction that you

**[00:32:47 – 00:32:50]** would find in CUDA, whereas that's not the case in Triton.

**[00:32:50 – 00:32:54]** So in Triton, all your functions are decorated with Triton.JIT,

**[00:32:54 – 00:32:57]** which means when you design your code, you can treat all

**[00:32:57 – 00:33:00]** your functions as device functions and then instantiate them

**[00:33:00 – 00:33:04]** as kernels as and when required, essentially at runtime.

**[00:33:04 – 00:33:07]** That leads to leaner code bases, and it's easier to

**[00:33:07 – 00:33:11]** design from that aspect.

**[00:33:11 – 00:33:17]** When it comes to the middle layer, Tile.ir has this view-based

**[00:33:17 – 00:33:22]** load in stores, which simplifies your kernels to a large extent.

**[00:33:22 – 00:33:25]** The memory consistency model that it uses is slightly different

**[00:33:25 – 00:33:27]** from what PTX uses.

**[00:33:27 – 00:33:30]** It's based on token ordering.

**[00:33:30 – 00:33:33]** As I mentioned, there's also this operation to extract

**[00:33:33 – 00:33:37]** a subtitle from a tile, which is This is not supported in Triton.

**[00:33:38 – 00:33:41]** When it comes to performance debugging, as it stands today,

**[00:33:41 – 00:33:44]** it's better in Triton because you can actually dump the

**[00:33:44 – 00:33:49]** IR and view all the optimization pipeline passes that were run

**[00:33:49 – 00:33:52]** and just god-world your way through it, and it's a really rewarding

**[00:33:53 – 00:33:55]** experience to figure out what.

**[00:33:55 – 00:34:03]** So that's not something that you get to see in QTile today,

**[00:34:03 – 00:34:06]** and that's a pain point.

**[00:34:06 – 00:34:10]** And some compiler directives that you use in Triton, like number of

**[00:34:10 – 00:34:15]** warps or number of pipeline stages, are not available yet in QTile.

**[00:34:16 – 00:34:18]** Let's look at some of the programming strengths.

**[00:34:18 – 00:34:22]** So if we look at block-based programming in general,

**[00:34:23 – 00:34:27]** It frees up space for the programmer to do higher-level

**[00:34:27 – 00:34:32]** algorithm decomposition and mapping it to parallel work, while the

**[00:34:32 – 00:34:36]** compiler takes care of lower-level hardware-aware intricacies.

**[00:34:36 – 00:34:40]** So my view on this is it gives you better signal-to-noise ratio.

**[00:34:40 – 00:34:43]** You focus more on the things that matter for performance,

**[00:34:43 – 00:34:47]** while not worrying about the lower-level details.

**[00:34:47 – 00:34:50]** And I work for a company who is this crazy about

**[00:34:50 – 00:34:51]** signal-to-noise ratio.

**[00:34:52 – 00:34:54]** We like this.

**[00:34:54 – 00:34:59]** Now, in CUDA tile, on top of this, you also have array-based indexing,

**[00:34:59 – 00:35:03]** which leads to more intuitive APIs.

**[00:35:03 – 00:35:06]** So if you noticed, the lines of code was consistently lower

**[00:35:06 – 00:35:10]** for QTile Python when compared to Triton, and a big reason

**[00:35:10 – 00:35:13]** for that was because we were able to use these view-based

**[00:35:13 – 00:35:17]** loads, or array-based indexing, which reduced the number of

**[00:35:17 – 00:35:19]** lines of code significantly.

**[00:35:19 – 00:35:21]** It also is less error-prone.

**[00:35:21 – 00:35:22]** You've got implicit.

**[00:35:22 – 00:35:29]** Bounce tech, implicit mapping to TMA, so it's quite useful there.

**[00:35:29 – 00:35:33]** On top of this, the fact that Tile IR is a stable virtual ISC with

**[00:35:33 – 00:35:37]** forward compatibility for future hardware generations is something

**[00:35:37 – 00:35:40]** that's very exciting for us, which means you don't have to worry

**[00:35:40 – 00:35:44]** about all the hardware features and redesigning your kernel

**[00:35:44 – 00:35:47]** each time a new feature drops.

**[00:35:47 – 00:35:48]** What are some limitations?

**[00:35:48 – 00:35:52]** Well, for block-based programming languages, operations on shifted

**[00:35:52 – 00:35:57]** loads, operations that have got shifted loads on tiles,

**[00:35:57 – 00:35:58]** they tend to perform poorly.

**[00:35:58 – 00:36:01]** This is because you want to have explicit control over

**[00:36:01 – 00:36:06]** shared memory or tensor memory to really get the performance,

**[00:36:06 – 00:36:08]** which is not available now.

**[00:36:08 – 00:36:11]** And when you choose to do ninja optimizations nonetheless,

**[00:36:11 – 00:36:17]** this involves tweaking the IR or falling back to CUDA, C++, or PTX.

**[00:36:17 – 00:36:22]** CUDA tile, well, some of these gaps were bridged in 13.2,

**[00:36:22 – 00:36:25]** I should admit that, but again, these relate to the fact that

**[00:36:25 – 00:36:28]** you don't have control over all the low-level details, which

**[00:36:28 – 00:36:31]** are necessary for performance.

**[00:36:31 – 00:36:34]** And finally, for performance debugging,

**[00:36:34 – 00:36:37]** Today, only power of two tile sizes are supported.

**[00:36:37 – 00:36:41]** It was good to hear from Bryce that that will soon go away.

**[00:36:41 – 00:36:45]** And we also don't have control, we don't get to see tile IR

**[00:36:45 – 00:36:47]** in the MLIR format.

**[00:36:47 – 00:36:51]** And the fact that the mapping from tile IR to SAS is opaque is also

**[00:36:51 – 00:36:54]** something that affects performance.

**[00:36:54 – 00:36:57]** So finally, I'll invite Pradeep back on stage to share the

**[00:36:57 – 00:36:59]** concluding remarks.

**[00:36:59 – 00:37:02]** So, just to quickly conclude, I think we spoke about the

**[00:37:02 – 00:37:04]** fact that, you know, what KLA does, you saw that we

**[00:37:04 – 00:37:09]** leverage AI and advancements in HPC to continue scaling Moore's Law,

**[00:37:09 – 00:37:14]** so better GPUs, faster processors, KLA has been at the heart of this.

**[00:37:14 – 00:37:17]** We found that through our HPC experience and having

**[00:37:17 – 00:37:20]** good hardware software co-design is really critical to achieve

**[00:37:20 – 00:37:22]** good execution time on accelerator.

**[00:37:22 – 00:37:24]** And our choice of compute and programming languages

**[00:37:25 – 00:37:29]** go hand in hand in identifying this, and as we found in this

**[00:37:29 – 00:37:32]** experience, block-based programming languages are turning out

**[00:37:32 – 00:37:34]** to be a very promising abstraction for us, and we are very excited

**[00:37:34 – 00:37:36]** about what it brings to the table.

**[00:37:36 – 00:37:39]** I think we still have some ways to go before it truly

**[00:37:39 – 00:37:41]** becomes an alternative to CUDA, but the fact that you

**[00:37:41 – 00:37:44]** can pick and choose when to do this and when to go to CUDA is something

**[00:37:44 – 00:37:45]** that is very exciting for us.

**[00:37:46 – 00:37:48]** That's all we've got for today, I think we still have a minute

**[00:37:48 – 00:37:50]** and a half for some questions, I'll be happy to answer any questions.

**[00:37:50 – 00:37:51]** Thank you.

**[00:37:52 – 00:37:54]** There were two dimensions that were missing for me,

**[00:37:54 – 00:37:59]** which was the hardware generation supported by Triton and Qtile,

**[00:37:59 – 00:38:02]** and then the amount of code that's available from the

**[00:38:02 – 00:38:05]** communities that are actually working and providing code and

**[00:38:05 – 00:38:07]** open source code on each platform.

**[00:38:07 – 00:38:09]** Could you rate these?

**[00:38:09 – 00:38:12]** Yeah, so let me take this latter question, right?

**[00:38:12 – 00:38:15]** So I think Kutile is relatively much newer, and Triton has

**[00:38:15 – 00:38:17]** been around for relatively longer.

**[00:38:17 – 00:38:19]** It's still short, but relatively longer.

**[00:38:19 – 00:38:22]** So definitely, right now in the community, there's more Triton code

**[00:38:22 – 00:38:25]** available, but with every release, we are seeing more developers

**[00:38:25 – 00:38:27]** taking on Kutile as well, right?

**[00:38:27 – 00:38:29]** So I think it's...

**[00:38:29 – 00:38:32]** Maybe Steven can say better about what the adoption has

**[00:38:32 – 00:38:34]** been, but we do see a lot more people playing with both

**[00:38:34 – 00:38:36]** of those paradigms, right?

**[00:38:36 – 00:38:38]** So that's one. I know it's not a direct answer to your question, but

**[00:38:39 – 00:38:40]** that's something that we've seen.

**[00:38:40 – 00:38:44]** I think the initial release at Blackwell, so our evaluation

**[00:38:44 – 00:38:45]** here was primarily on Blackwell.

**[00:38:45 – 00:38:49]** I know the 13.2 added support for both Ada and for Ampere.

**[00:38:49 – 00:38:51]** We've not gotten a chance to try that out.

**[00:38:51 – 00:38:53]** But I think on the roadmap that I saw on Stephen's slide

**[00:38:53 – 00:38:56]** today, it's, you know, more support is coming up.

**[00:38:56 – 00:38:59]** The initial focus, of course, was on Blackwell, but not all the

**[00:38:59 – 00:39:00]** features of Blackwell got lit up.

**[00:39:00 – 00:39:03]** And this is developing with every CUDA release.

**[00:39:03 – 00:39:05]** So right now, that's why you'll see all our evaluation, in

**[00:39:05 – 00:39:06]** fact, was only on the 1590.

**[00:39:06 – 00:39:10]** It's not even the RTX Pro 6000, but more support is coming.

**[00:39:10 – 00:39:12]** Hi, thank you for the talk.

**[00:39:12 – 00:39:15]** I have a question about the shared memory and how it's

**[00:39:15 – 00:39:19]** being used automatically by the Gutao compiler.

**[00:39:19 – 00:39:24]** So from your experience, how well did the compiler manage the

**[00:39:24 – 00:39:28]** trade-off between the low latency of putting things into shared

**[00:39:28 – 00:39:32]** memory versus the thread occupancy that you might lose because

**[00:39:32 – 00:39:37]** you're putting too much things into the shared memory to fit on the SM?

**[00:39:37 – 00:39:41]** All right, so I should confess that I don't have a good answer

**[00:39:41 – 00:39:45]** to that, because it's quite non-trivial to figure out

**[00:39:45 – 00:39:46]** when shared memory is used.

**[00:39:47 – 00:39:49]** And that's one of the things where Triton really shines, right?

**[00:39:49 – 00:39:52]** You can actually see what happens after the shared memory

**[00:39:52 – 00:39:57]** pass in Triton and see the decisions that it took.

**[00:39:57 – 00:40:01]** But regardless, if you look at inside compute reports, you can see

**[00:40:01 – 00:40:03]** that shared memory is allocated.

**[00:40:03 – 00:40:07]** Very easy to pinpoint which part of the code it was being used.

**[00:40:07 – 00:40:11]** But yes, that said, from a performance angle, I think

**[00:40:11 – 00:40:14]** the numbers are testament to the fact that it does do

**[00:40:14 – 00:40:17]** a good job on balancing the occupancy and resource utilization.

