crossbind
GitHub

Zstandard for Linux

v1.5.7Linux

Zstandard 1.5.7 for native Node.js addons on Linux, precompiled for x64 and arm64, glibc 2.28 or later as @crossbind/port-zstd-linux.

npm install @crossbind/port-zstd-linux@beta

Install

shell
npm install @crossbind/port-zstd-linux@beta
npm install --save-dev crossbind@beta @crossbind/core-embind-napi@beta
crossbind.config.js
import zstdLinux from '@crossbind/port-zstd-linux/crossbind.config.js';
 
export default {
dependencies: [zstdLinux],
paths: { config: import.meta.url },
};
shell
npx crossbind build -p linux

The addons and their loader land in dist; the Node.js playbook has the whole flow.

Usage

The examples the WebAssembly page runs, as Linux compiles them: the same headers and the same calls. They are checked on the WebAssembly build.

Compress and decompress a buffer

The two calls most zstd code makes: ZSTD_compress, and ZSTD_decompress with the original size read back from the frame.

src/native/zstd_codec.h
#pragma once
 
#include <zstd.h>
 
#include <stdexcept>
#include <string>
 
// One-shot Zstandard. Bytes cross the binding as a byte string: one UTF-16 code unit (0-255) per byte.
class Zstd {
public:
static std::string version() { return ZSTD_versionString(); }
 
static std::u16string compress(const std::string& text, int level) {
std::string out(ZSTD_compressBound(text.size()), '\0');
const size_t size = ZSTD_compress(&out[0], out.size(), text.data(), text.size(), level);
if (ZSTD_isError(size)) throw std::runtime_error(ZSTD_getErrorName(size));
std::u16string bytes(size, u'\0');
for (size_t i = 0; i < size; ++i) bytes[i] = static_cast<unsigned char>(out[i]);
return bytes;
}
 
static std::string decompress(const std::u16string& bytes) {
std::string in(bytes.size(), '\0');
for (size_t i = 0; i < bytes.size(); ++i) {
if (bytes[i] > 0xFF) throw std::invalid_argument("not a byte string");
in[i] = static_cast<char>(bytes[i]);
}
const unsigned long long size = ZSTD_getFrameContentSize(in.data(), in.size());
if (size == ZSTD_CONTENTSIZE_ERROR) throw std::runtime_error("not a zstd frame");
if (size == ZSTD_CONTENTSIZE_UNKNOWN) throw std::runtime_error("size not stored in the frame; use streaming");
if (size > (256u << 20)) throw std::runtime_error("refusing to allocate more than 256 MiB");
std::string out(static_cast<size_t>(size), '\0');
const size_t got = ZSTD_decompress(&out[0], out.size(), in.data(), in.size());
if (ZSTD_isError(got)) throw std::runtime_error(ZSTD_getErrorName(got));
out.resize(got);
return out;
}
};
main.js
import { initNative, Zstd } from './native/zstd_codec.h';
 
await initNative();
const text = 'crossbind '.repeat(100);
const frame = await Zstd.compress(text, 19);
const bytes = Uint8Array.from(frame, (c) => c.charCodeAt(0));
console.log(await Zstd.version(), bytes.length, [...bytes.slice(0, 4)].map((b) => b.toString(16)).join(' '));
console.log((await Zstd.decompress(frame)) === text);
PRINTS
1.5.7 27 28 b5 2f fd
true

Choose a level, a checksum and a window

A context set up with ZSTD_CCtx_setParameter decides what every frame it writes looks like; ZSTD_getFrameHeader reads the choices back.

src/native/zstd_frame.h
#pragma once
 
// ZSTD_getFrameHeader is in zstd's static API; safe with this statically linked, pinned library.
#define ZSTD_STATIC_LINKING_ONLY
#include <zstd.h>
 
#include <memory>
#include <stdexcept>
#include <string>
 
// A compression context with explicit parameters, and a reader for what they put in the frame header.
class ZstdFrame {
public:
// A windowLog of 0 keeps the level's default window.
static std::u16string compress(const std::string& text, int level, bool checksum, int windowLog) {
std::unique_ptr<ZSTD_CCtx, size_t (*)(ZSTD_CCtx*)> cctx(ZSTD_createCCtx(), ZSTD_freeCCtx);
check(ZSTD_CCtx_setParameter(cctx.get(), ZSTD_c_compressionLevel, level));
check(ZSTD_CCtx_setParameter(cctx.get(), ZSTD_c_checksumFlag, checksum ? 1 : 0));
if (windowLog) check(ZSTD_CCtx_setParameter(cctx.get(), ZSTD_c_windowLog, windowLog));
std::string out(ZSTD_compressBound(text.size()), '\0');
out.resize(check(ZSTD_compress2(cctx.get(), &out[0], out.size(), text.data(), text.size())));
std::u16string bytes(out.size(), u'\0');
for (size_t i = 0; i < out.size(); ++i) bytes[i] = static_cast<unsigned char>(out[i]);
return bytes;
}
 
static std::string header(const std::u16string& frame) {
std::string head;
for (size_t i = 0; i < frame.size() && i < ZSTD_FRAMEHEADERSIZE_MAX; ++i) head += static_cast<char>(frame[i]);
ZSTD_FrameHeader info;
const size_t status = ZSTD_getFrameHeader(&info, head.data(), head.size());
if (ZSTD_isError(status) || status > 0) throw std::runtime_error("not a zstd frame header");
const std::string content = info.frameContentSize == ZSTD_CONTENTSIZE_UNKNOWN ? "not stored" : std::to_string(info.frameContentSize) + " B";
return "content " + content + ", window " + std::to_string(info.windowSize) + " B, checksum " + (info.checksumFlag ? "yes" : "no");
}
 
private:
static size_t check(size_t code) {
if (ZSTD_isError(code)) throw std::runtime_error(ZSTD_getErrorName(code));
return code;
}
};
main.js
import { initNative, ZstdFrame } from './native/zstd_frame.h';
 
await initNative();
let seed = 7;
const random = (n) => (seed = (seed * 48271) % 2147483647) % n;
const text = Array.from({ length: 400 }, (_, i) => `{"id":${i},"user":"user${random(5000)}","score":${random(1000)}}`).join('\n');
for (const level of [3, 9, 19]) {
const frame = await ZstdFrame.compress(text, level, false, 0);
console.log(`level ${level}: ${text.length} B -> ${frame.length} B`);
}
const small = await ZstdFrame.compress(text, 19, true, 10);
console.log(`level 19, 1 KiB window, checksum: ${small.length} B`);
console.log(await ZstdFrame.header(small));
PRINTS
level 3: 16160 B -> 3065 B
level 9: 16160 B -> 2749 B
level 19: 16160 B -> 2219 B
level 19, 1 KiB window, checksum: 2469 B
content 16160 B, window 1024 B, checksum yes

Compress small messages with a dictionary

Train once with ZDICT_trainFromBuffer, digest it once with ZSTD_createCDict and ZSTD_createDDict, then compress every message against it.

src/native/zstd_dictionary.h
#pragma once
 
#include <zdict.h>
#include <zstd.h>
 
#include <memory>
#include <stdexcept>
#include <string>
#include <vector>
 
// Dictionary compression for small messages: train once on samples, prepare the dictionary once,
// then compress and decompress each message against it.
class ZstdDictionary {
public:
// Trains on newline-separated samples; returns at most `capacity` bytes of dictionary.
static std::u16string train(const std::string& samples, int capacity) {
std::string joined;
std::vector<size_t> sizes;
for (size_t start = 0; start < samples.size();) {
size_t end = samples.find('\n', start);
if (end == std::string::npos) end = samples.size();
joined.append(samples, start, end - start);
sizes.push_back(end - start);
start = end + 1;
}
std::string dictionary(static_cast<size_t>(capacity), '\0');
const size_t size = ZDICT_trainFromBuffer(&dictionary[0], dictionary.size(), joined.data(), sizes.data(), static_cast<unsigned>(sizes.size()));
if (ZDICT_isError(size)) throw std::runtime_error(ZDICT_getErrorName(size));
dictionary.resize(size);
return toUnits(dictionary);
}
 
// Digests the dictionary once for compression and once for decompression.
ZstdDictionary(const std::u16string& dictionary, int level)
: compressDictionary(nullptr, ZSTD_freeCDict), decompressDictionary(nullptr, ZSTD_freeDDict) {
const std::string bytes = fromUnits(dictionary);
compressDictionary.reset(ZSTD_createCDict(bytes.data(), bytes.size(), level));
decompressDictionary.reset(ZSTD_createDDict(bytes.data(), bytes.size()));
if (!compressDictionary || !decompressDictionary) throw std::runtime_error("not a usable dictionary");
}
 
std::u16string compress(const std::string& message) const {
std::unique_ptr<ZSTD_CCtx, size_t (*)(ZSTD_CCtx*)> cctx(ZSTD_createCCtx(), ZSTD_freeCCtx);
std::string out(ZSTD_compressBound(message.size()), '\0');
out.resize(check(ZSTD_compress_usingCDict(cctx.get(), &out[0], out.size(), message.data(), message.size(), compressDictionary.get())));
return toUnits(out);
}
 
std::string decompress(const std::u16string& frame) const {
const std::string in = fromUnits(frame);
const unsigned long long size = ZSTD_getFrameContentSize(in.data(), in.size());
if (size == ZSTD_CONTENTSIZE_ERROR || size == ZSTD_CONTENTSIZE_UNKNOWN) throw std::runtime_error("not a zstd frame with its size");
if (size > (1u << 20)) throw std::runtime_error("a message above 1 MiB is not a small message");
std::unique_ptr<ZSTD_DCtx, size_t (*)(ZSTD_DCtx*)> dctx(ZSTD_createDCtx(), ZSTD_freeDCtx);
std::string out(static_cast<size_t>(size), '\0');
out.resize(check(ZSTD_decompress_usingDDict(dctx.get(), &out[0], out.size(), in.data(), in.size(), decompressDictionary.get())));
return out;
}
 
private:
static size_t check(size_t code) {
if (ZSTD_isError(code)) throw std::runtime_error(ZSTD_getErrorName(code));
return code;
}
 
static std::u16string toUnits(const std::string& data) {
std::u16string units(data.size(), u'\0');
for (size_t i = 0; i < data.size(); ++i) units[i] = static_cast<unsigned char>(data[i]);
return units;
}
 
static std::string fromUnits(const std::u16string& units) {
std::string data(units.size(), '\0');
for (size_t i = 0; i < units.size(); ++i) {
if (units[i] > 0xFF) throw std::invalid_argument("not a byte string");
data[i] = static_cast<char>(units[i]);
}
return data;
}
 
std::unique_ptr<ZSTD_CDict, size_t (*)(ZSTD_CDict*)> compressDictionary;
std::unique_ptr<ZSTD_DDict, size_t (*)(ZSTD_DDict*)> decompressDictionary;
};
main.js
import { initNative, ZstdDictionary } from './native/zstd_dictionary.h';
import { Zstd } from './native/zstd_codec.h';
 
await initNative();
const event = (i) => `{"event":"click","user":${1000 + ((i * 37) % 900)},"page":"/products/${i % 12}","ms":${(i * 7919) % 400}}`;
const samples = Array.from({ length: 4000 }, (_, i) => event(i)).join('\n');
const dictionary = await ZstdDictionary.train(samples, 2048);
const codec = await new ZstdDictionary(dictionary, 3);
 
const message = event(4321);
const alone = await Zstd.compress(message, 3); // the one-shot wrapper from the first example
const frame = await codec.compress(message);
console.log(`dictionary: ${dictionary.length} B`);
console.log(`${message.length} B message: ${alone.length} B alone, ${frame.length} B with the dictionary`);
console.log((await codec.decompress(frame)) === message);
PRINTS
dictionary: 2048 B
59 B message: 68 B alone, 29 B with the dictionary
true

"Stream a file through zstd" writes its input with m.FS, which Linux does not have; the C++ takes paths, so it works unchanged on files in the app's storage. It runs on the WebAssembly page.

What is different on Linux

  • crossbind build -p linux links one .node addon per architecture into dist, next to a loader, dist/<name>.native.cjs, that require and import both load. A plain crossbind build skips it.
  • await initNative() once, then call the classes: calls are synchronous, and no Worker is involved.
  • The library and the C++ runtime are linked into the addon statically. It runs on glibc 2.28 or later (RHEL 8, Debian 10, Ubuntu 20.04 and newer), not on musl distributions such as Alpine.
  • There is no m.FS: the C++ reads real paths, and data such as GDAL_DATA or proj.db is copied to dist/data.
  • The build runs in Docker on any host, a Mac included. worker_threads are not supported yet.

Other platforms

Facts on this page come from the port manifests in the repository and from what npm served on beta when the site was built. See the Libraries guide for the full consumer flow.

MORE LIBRARIES
cURLExpatGDALGEOSGeoTIFFiconvLERClibjpeg-turbolibTIFFOpenSSLPROJSpatiaLiteSQLiteWebPzlib
Type to search every guide page and section.
↑↓ navigate↵ openesc close