bcn is a native Go codec for BCn GPU texture compression (S3TC and BPTC).
It encodes and decodes texture blocks, works directly with Go images,
and reads and writes DDS and KTX v1 containers.
Use it to turn image.Image values into GPU-ready textures,
inspect or decode existing DDS/KTX assets,
and build 2D or cubemap textures with mipmaps.
The hot paths use AVX2/SSE2 kernels on amd64 when available;
-tags purego builds the same API without assembly or cgo.
- BC1 / DXT1, BC2 / DXT3, BC3 / DXT5, BC4 / BC5 (UNORM and SNORM), BC6H / BPTC-HDR, BC7 / BPTC encode/decode
- DDS read/write (2D + cubemap, mipmaps, uncompressed RGBA8 / BGRA8 / BGRX8 / A8 / R8 / RG8/ R8S / RG8S / RGB10A2 / RGB565 / RGBA5551 / RGBA4444 / RGB8 / BGR8)
- KTX v1 read/write (2D + cubemap, mipmaps, uncompressed RGBA8 / BGRA8 / A8 / R8 / RG8 / R8S / RG8S / RGB10A2 / RGB565 / RGBA5551 / RGBA4444 / RGB8 / BGR8)
- Mipmap generation with optional sRGB-aware downscale
- Quality levels (1..10) with least-squares endpoint refit
and refinement overrides (
Refinement) - Parallel encoding control via
EncodeOptions.Workers(0=auto, 1=off) - Parallel decoding control via
DecodeOptions.Workers(0=auto, 1=off)
BC4/BC5 signed normalized variants use FormatBC4S / FormatBC5S.
Their NRGBA input and output map 0..255 to -1..1.
Note
For large images or one‑by‑one encoding, use internal parallelism (default).
For batch/many small files, parallelize across images in your own code
and keep Workers=1 here.
Workers=0 uses GOMAXPROCS (Go scheduler's CPU limit).
img, _, _ := image.Decode(in)
opts := &bcn.EncodeOptions{
QualityLevel: bcn.QualityLevelBalanced,
GenerateMipmaps: true,
UseSRGB: true,
}
dds, err := bcn.EncodeDDSWithOptions([]image.Image{img}, bcn.FormatBC3, opts)
if err != nil {
/* handle */
}
_ = dds.Write(out)dds, err := bcn.ReadDDS(in)
if err != nil {
/* handle */
}
img, err := bcn.DecodeImage(dds.Faces[0].Mipmaps[0], dds.Width, dds.Height, dds.Format)
if err != nil {
/* handle */
}
_ = png.Encode(out, img)hdr, dx10, err := bcn.ReadDDSHeader(in)
if err != nil {
/* handle */
}
_ = hdr
_ = dx10ktx, err := bcn.EncodeKTXWithOptions([]image.Image{img}, bcn.FormatBC5, &bcn.EncodeOptions{QualityLevel: bcn.QualityLevelFast})
if err != nil {
/* handle */
}
_ = ktx.Write(out)Import the subpackages to register DDS and KTX with the image package;
then image.Decode and image.DecodeConfig work as usual:
import (
_ "github.com/woozymasta/bcn/dds"
_ "github.com/woozymasta/bcn/ktx"
)
// ...
img, _, _ := image.Decode(f) // decodes first face/mip to NRGBA
cfg, _, _ := image.DecodeConfig(f) // width, height only- KTX v1 arrays and 3D textures are not supported.
- DDS DX10 supports BC1–BC7, RGBA/BGRA/BGRX, A8, R8/RG8 (UNORM and SNORM), RGB10A2, RGB565, RGBA5551, RGBA4444, RGB8, and BGR8; writing uses legacy FourCC where available.
- BC4 uses red channel; BC5 uses red/green.
- DDS BGRA is converted to RGBA on decode; BGRX always decodes with alpha
255. R8 decodes asR,R,R,255; RG8 asR,G,0,255. Signed R8S/RG8S use the same0..255to-1..1mapping as BC4S/BC5S. A8 decodes as0,0,0,A. - RGB10A2 is packed
R:10,G:10,B:10,A:2UNORM; conversion to/from NRGBA uses nearest rounding. - RGB565, RGBA5551, and RGBA4444 use nearest rounding when converting to/from NRGBA.
RefinementoverridesQualityLevelwhen set.- Quality levels above 1 polish endpoints
with an iterated least-squares refit on top of the grid search
(higher quality, some extra encode cost; decode is unaffected).
Disable or tune it viaRefinement.LSQIters(0= off,nil= quality default,N= iterations); setColorTries: 0withLSQIters > 0for a cheap LSQ-only refine.
On amd64 the hot encode/decode paths use AVX2/SSE2 assembly kernels
(in the internal/simd package, generated with avo)
selected at runtime via golang.org/x/sys/cpu;
a portable pure-Go fallback handles every other platform
and any block the kernels do not cover
(e.g. edge blocks when width or height is not a multiple of 4).
The two paths are byte-exact -
validated by exhaustive, randomized and fuzz equivalence tests.
- AVX2 kernels need AVX2 (decode of BC1 needs AVX2; BC2/BC3/BC4/BC5 decode needs AVX2+BMI2). Without them the pure-Go path runs.
BCN_PUREGO=1in the environment forces the pure-Go path at runtime; building with-tags puregoexcludes the assembly entirely.- The avo generator lives in its own build-time-only module
(
internal/simd/asmgen), so consumers never pull avo into their module graph; the only runtime dependency isgolang.org/x/sys. - Regenerate kernels after editing
internal/simd/asmgenwithmake generate;make generate-check(part of CI) verifies the committed.sis up to date. - Set
GOAMD64=v2/v3to let the Go compiler also vectorize the fallback and container helpers; it does not affect the hand-written kernels.
Single-thread, Ryzen 9 5950X, Go 1.26, 512x512, throughput over input bytes (RGBA for LDR formats, RGB float16 for BC6H; higher is better):
| Format | fast, MB/s | balanced, MB/s | best, MB/s | decode, MB/s |
|---|---|---|---|---|
| BC1 | ~830 | ~35 | ~12 | ~920 |
| BC2 | ~635 | ~33 | ~12 | ~1715 |
| BC3 | ~370 | ~31 | ~11 | ~1670 |
| BC4 | ~480 | ~74 | ~23 | ~1530 |
| BC5 | ~255 | ~38 | ~12 | ~2490 |
| BC6H | ~150 | ~8 | ~2 | ~150 |
| BC7 | ~55 | ~0.9 | ~0.5 | ~66 |
Multi-thread, Workers=auto (GOMAXPROCS=32), 512x512,
encode throughput over input bytes (higher is better):
| Format | fast, MB/s | balanced, MB/s | best, MB/s |
|---|---|---|---|
| BC1 | ~5,000 | ~380 | ~144 |
| BC3 | ~2,620 | ~370 | ~130 |
| BC6H | ~1,350 | ~90 | ~30 |
| BC7 | ~540 | ~13 | ~7 |
Fast/Balanced/Best correspond to
QualityLevelFast, QualityLevelBalanced, QualityLevelBest.
For batch/many small files,
parallelize across images in your own code and keep Workers=1;
see EncodeOptions.Workers.