TinyBCDec is a tiny library for decoding block compressed texture formats. It's written in pure Java (11+), without any dependencies, with a focus on speed and accuracy.
Currently, the following formats are supported:
- BC1 (DXT1)
- BC2 (DXT3)
- BC3 (DXT5)
- BC4 (ATI1)
- BC5 (ATI2)
- BC6H
- BC7
Every format except BC6H is decoded as BGRA, 4 bytes per pixel. Read as little endian integers, that is AWT's
TYPE_INT_ARGB layout, so the result can go straight into a BufferedImage without a channel swap.
- BC1, BC2, BC3 and BC7 fill all four channels.
- BC4 is a single channel, expanded to gray: blue, green and red all get the value, alpha is 255.
- BC5 is two channels: red and green carry the data, blue is 0 and alpha is 255.
- BC6H is the exception: it is decoded as RGBA in little endian half-float (8 bytes per pixel), in that channel order rather than BGRA, with alpha set to 1.0.
Every format also has a float variant, from the factory methods ending in Float. These decode to RGBA
single-precision floats, 16 bytes per pixel. Mind the channel order: float output is RGBA, like BC6H, and not BGRA
like the byte output.
- BC1 through BC5 and BC7 carry the same channels as their byte output, normalized to the range 0 to 1.
- BC6H is already floating point, so
bc6hFloatonly widens each half-float channel to a full float.
Using the library is straightforward, all you need is the compressed data, and the format of the data.
The following features are present:
- Partial decodes: The width and height do not need to be a multiple of the block size (4 in this case). The output width and height can be smaller than the input, and the library will handle this.
- BC6H: BC6H is decoded to a little endian half-float RGBA buffer, 8 bytes per pixel.
- Float output: every format has a matching
Floatfactory method, decoding straight to RGBA floats at 16 bytes per pixel, so there is no conversion pass afterwards.
The library provides the BlockDecoder class, which can be used to decode. A new instance is created by one of the
static factory methods. You can let the library create a new buffer, or pass an existing one to save allocations.
After construction, decoders produce no heap allocations during decoding. Internal working state is pre-allocated and reused across calls. Note that this also means instances are not thread-safe: each thread must use its own decoder instance.
Both the source and the destination are a ByteBuffer. They are read and written by absolute index, starting at the
buffer's current position, so decoding never advances a position. Block data is little endian, so both buffers are
switched to ByteOrder.LITTLE_ENDIAN while decoding and set back to their original order afterwards.
Given src is the compressed data, positioned at the start of the blocks, the following snippet decodes a BC1 texture
into a newly allocated buffer.
import be.twofold.tinybcdec.*;
BlockDecoder decoder = BlockDecoder.bc1(true);
ByteBuffer result = decoder.decode(src, 256, 256);To decode into a buffer you already have, pass it along with its dimensions. There is no return value.
decoder.decode(src, 256, 256, dst, 256, 256);The destination may be smaller than the source, in which case the image is cropped to the top left. It may not be larger, as there would be nothing to fill the remainder with.
decoder.decode(src, 256, 256, dst, 64, 64); // the top left 64x64 pixelsTo decode somewhere other than the top left, add the source and destination coordinates. The region decoded is whatever is left of the destination, which makes this the form to use when the destination is already the size of the region you want, such as a single tile.
decoder.decode(src, 256, 256, tile, 64, 64, 128, 64, 0, 0); // the 64x64 tile at (128, 64)The general form takes the region size explicitly, and is the only way to decode into part of a larger destination.
decoder.decode(
src, 256,256, // source buffer and its dimensions
dst, 256,256, // destination buffer and its dimensions
128, 64, 0, 0, // decode from (128, 64) to (0, 0)
64,64 // decoding a 64x64 region
);Both buffers must hold their entire image, not just the region, since the dimensions are what determine where a row
starts. encodedByteSize(width, height) and decodedByteSize(width, height) return the sizes required.
The BGRA output maps directly onto the common image types, so these snippets need no per-pixel work. They apply to the byte output only, not to BC6H or the float variants, whose floating point output does not fit either.
For an AWT BufferedImage, the bytes are TYPE_INT_ARGB when read as little endian integers:
ByteBuffer bgra = decoder.decode(src, width, height);
BufferedImage image = new BufferedImage(width, height, BufferedImage.TYPE_INT_ARGB);
int[] pixels = ((DataBufferInt) image.getRaster().getDataBuffer()).getData();
bgra.order(ByteOrder.LITTLE_ENDIAN).asIntBuffer().get(pixels);For a JavaFX Image, the bytes are straight-alpha BGRA, which is PixelFormat.getByteBgraInstance():
ByteBuffer bgra = decoder.decode(src, width, height);
WritableImage image = new WritableImage(width, height);
image.getPixelWriter().setPixels(
0,0,width, height,
PixelFormat.getByteBgraInstance(),bgra,width *4);The float output takes a little more work, as there is no ready made image type for it. This is BC6H, but any of the
Float decoders works the same way, since they all produce four floats per pixel:
ByteBuffer rgba = BlockDecoder.bc6hFloat(false).decode(src, width, height);
float[] pixels = new float[width * height * 4];
rgba.order(ByteOrder.LITTLE_ENDIAN).asFloatBuffer().get(pixels);
// AWT only deals in the 0 to 1 range, so anything outside it has to go before drawing.
for (int i = 0; i < pixels.length; i++) {
pixels[i] = Math.max(0.0f, Math.min(1.0f, pixels[i])); // Or Math.clamp
}
ColorModel colorModel = new ComponentColorModel(
ColorSpace.getInstance(ColorSpace.CS_LINEAR_RGB), // or CS_sRGB
true, false, Transparency.TRANSLUCENT, DataBuffer.TYPE_FLOAT);
WritableRaster raster = Raster.createWritableRaster(
new PixelInterleavedSampleModel(
DataBuffer.TYPE_FLOAT, width, height, 4, width * 4, new int[]{0, 1, 2, 3}),
new DataBufferFloat(pixels, pixels.length), null);
BufferedImage image = new BufferedImage(colorModel, raster, false, null);That clamp is not optional. BC6H is a high dynamic range format, so channels regularly go above 1.0, and a linear
ComponentColorModel looks colors up in a table sized for the 0 to 1 range. Drawing an image with values outside it
throws an ArrayIndexOutOfBoundsException, rather than clamping. Clamping flattens everything brighter than white, so
tone map instead if the highlights matter.
Version 2.0 changed both the API and the output, so it is not a drop-in replacement.
- Buffers are a
ByteBufferinstead of abyte[], and thesrcPosanddstPosparameters are gone. Reading and writing start at the buffer's own position, sodecode(data, offset, width, height)becomesdecode(buffer.position(offset), width, height). - The 8 bit output is BGRA, where it used to be RGBA. BC6H is unaffected, its half-float output was and stays RGBA. So if your red and blue channels seem swapped, this is why.
byteSizeis split intoencodedByteSizeanddecodedByteSize, for the source and the destination respectively.decodeBlockis no longer public. Never really useful, and no validation.
New in this version is the float output described above, and decoding is quite a bit faster: BC5 by about 30% and BC7 by about 20%.
I've done some performance testing, and the library is quite fast.
Some quick benchmarks, decoding a 256x256 texture on a Ryzen 5950X with Oracle Java 25 (MP/s stands for megapixels per second):
- BC1: ~1000MP/s
- BC2: ~650MP/s
- BC3: ~550MP/s
- BC4: ~1100MP/s
- BC5: ~550MP/s
- BC6: ~150MP/s
- BC7: ~200MP/s
These numbers are just an estimate, and can vary depending on the hardware and the JVM.
To give you an idea, this means about 17 ms to decode a 4096x4096 texture in BC1, and about 110 ms for BC6 or 85 ms for BC7.
A final note on accuracy, the library is tested against the output of DirectXTex. I generated images, encoded them, decoded them again and compared the output. The output is identical, except for signed BC4 and BC5, which are different by a small amount. This is due to the way they are handled in DirectXTex, which is different from the unsigned code path. As far as I can test, the values are correct on my side. A bug has been filed.
This is done by using fixed point arithmetic for BC1 through 5. Check out GenerateRescale.java for how these scale factors are calculated. My unsigned results all line up with what bcdec does.