Version: 1.3.0 | Last Updated: 2025-11-08
Complete API reference for all DataFrame operations available in the Node.js/TypeScript environment.
- Initialization
- DataFrame Creation
- CSV Operations
- DataFrame Utilities
- Column Operations
- Row Operations
- Missing Data
- String Operations
- Numeric Operations
- Aggregations
- Advanced Aggregations
- Window Operations
- Sorting
- Joins
- Grouping
- Reshape Operations
- Apache Arrow
- Lazy Evaluation
- Memory Management
Initialize the Rozes WASM module.
Parameters:
wasmPath(optional): Path torozes.wasmfile. Auto-detected in most environments.
Returns: Promise - Initialized Rozes instance
Example:
import { Rozes } from 'rozes';
// Browser
const rozes = await Rozes.init();
// Node.js (auto-detection)
const rozes = await Rozes.init();
// Node.js (explicit path)
const rozes = await Rozes.init('./node_modules/rozes/zig-out/bin/rozes.wasm');
const DataFrame = rozes.DataFrame;Create DataFrame from CSV string.
Parameters:
csvString: string- CSV data as stringoptions(optional):delimiter?: string- Column delimiter (default:,)hasHeader?: boolean- Has header row (default:true)quoteChar?: string- Quote character (default:")escapeChar?: string- Escape character (default:")
Returns: DataFrame
Example:
const csv = `name,age,city
Alice,30,NYC
Bob,25,LA`;
const df = DataFrame.fromCSV(csv);
df.show();
// name age city
// 0 Alice 30 NYC
// 1 Bob 25 LACreate DataFrame from column definitions and data arrays.
Parameters:
columns: ColumnDef[]- Array of column definitionsdata: any[][]- 2D array of row data
Example:
const columns = [
{ name: 'name', type: 'String' },
{ name: 'age', type: 'Int64' }
];
const data = [
['Alice', 30],
['Bob', 25]
];
const df = DataFrame.create(columns, data);Export DataFrame to CSV string.
Parameters:
options(optional):includeHeaders?: boolean- Include header row (default:true)delimiter?: string- Column delimiter (default:,)lineEnding?: string- Line ending (default:\n)
Returns: string - CSV formatted data
Example:
const csv = df.toCSV();
console.log(csv);
// name,age,city
// Alice,30,NYC
// Bob,25,LA
// Custom delimiter
const tsv = df.toCSV({ delimiter: '\t' });
// Without headers
const csvNoHeader = df.toCSV({ includeHeaders: false });Export DataFrame to CSV file.
Parameters:
path: string- Output file pathoptions(optional): Same astoCSV()
Example:
import fs from 'fs';
// Method 1: Using toCSV() + fs.writeFileSync
const csvData = df.toCSV();
fs.writeFileSync('output.csv', csvData, 'utf8');
// Method 2: Direct export (if implemented)
// df.toCSVFile('output.csv');Get DataFrame dimensions.
Returns: Object with rows and cols properties
Example:
console.log(df.shape);
// { rows: 100, cols: 5 }
console.log(`DataFrame has ${df.shape.rows} rows`);Get column names.
Returns: Array of column name strings
Example:
console.log(df.columns);
// ['name', 'age', 'city', 'score']
// Check if column exists
if (df.columns.includes('age')) {
// ...
}Get column data types.
Returns: Object mapping column names to types
Example:
console.log(df.dtypes);
// { name: 'String', age: 'Int64', score: 'Float64' }Display DataFrame (console output).
Parameters:
n?: number- Number of rows to show (default: all)
Example:
df.show(); // Show all rows
df.show(10); // Show first 10 rowsGet first N rows.
Parameters:
n?: number- Number of rows (default: 5)
Returns: New DataFrame with first N rows
Example:
const top5 = df.head(); // First 5 rows
const top10 = df.head(10); // First 10 rowsGet last N rows.
Parameters:
n?: number- Number of rows (default: 5)
Returns: New DataFrame with last N rows
Example:
const bottom5 = df.tail(); // Last 5 rows
const bottom10 = df.tail(10); // Last 10 rowsDrop columns by name.
Parameters:
columnNames: string | string[]- Column name(s) to drop
Returns: New DataFrame without specified columns
Example:
// Drop single column
const df2 = df.drop('age');
// Drop multiple columns
const df3 = df.drop(['age', 'city']);Rename a column.
Parameters:
oldName: string- Current column namenewName: string- New column name
Returns: New DataFrame with renamed column
Example:
const df2 = df.rename('age', 'years');
console.log(df2.columns);
// ['name', 'years', 'city']Get unique values in column.
Parameters:
columnName: string- Column name
Returns: Array of unique values
Example:
const cities = df.unique('city');
console.log(cities);
// ['NYC', 'LA', 'Chicago', 'Boston']Remove duplicate rows.
Parameters:
columnNames?: string[]- Columns to check for duplicates (default: all)
Returns: New DataFrame without duplicates
Example:
// Drop rows with duplicate values in all columns
const df2 = df.dropDuplicates();
// Drop rows with duplicate city values
const df3 = df.dropDuplicates(['city']);
// Drop based on multiple columns
const df4 = df.dropDuplicates(['name', 'age']);Statistical summary of numeric columns.
Parameters:
columnName?: string- Specific column (default: all numeric)
Returns: DataFrame with statistics (count, mean, std, min, max, etc.)
Example:
// Describe all numeric columns
const stats = df.describe();
stats.show();
// Describe specific column
const ageStats = df.describe('age');Random sample of rows.
Parameters:
n: number- Number of rows to sampleseed?: number- Random seed for reproducibility
Returns: DataFrame with N randomly sampled rows
Example:
// Random 10 rows
const sample = df.sample(10);
// Reproducible sample
const sample2 = df.sample(10, 42);Select specific columns.
Parameters:
columnNames: string[]- Column names to select
Returns: New DataFrame with only selected columns
Example:
const subset = df.select(['name', 'age']);
subset.show();
// name age
// 0 Alice 30
// 1 Bob 25Get column data.
Parameters:
columnName: string- Column name
Returns: Column object or null if not found
Example:
const ageCol = df.column('age');
if (ageCol) {
console.log(ageCol.data); // TypedArray or Array
console.log(ageCol.type); // 'Int64', 'Float64', etc.
}Add or replace column.
Parameters:
columnName: string- New column namevalues: any[]- Column values (must match row count)
Returns: New DataFrame with added/replaced column
Example:
// Add new column
const df2 = df.withColumn('category', ['A', 'B', 'A', 'B']);
// Replace existing column
const df3 = df.withColumn('age', [31, 26, 36, 29]);
// Computed column (from existing data)
const ages = df.column('age').data;
const doubledAges = Array.from(ages).map(a => a * 2);
const df4 = df.withColumn('age_doubled', doubledAges);Filter rows by condition.
Parameters:
predicate: (row: RowRef) => boolean- Filter function
Returns: New DataFrame with rows matching predicate
Example:
// Filter by age
const adults = df.filter(row => row.get('age') >= 30);
// Filter by string match
const nycOnly = df.filter(row => row.get('city') === 'NYC');
// Multiple conditions
const filtered = df.filter(row =>
row.get('age') > 25 && row.get('score') > 80
);Get rows by index range.
Parameters:
start: number- Start index (inclusive)end: number- End index (exclusive)
Returns: New DataFrame with sliced rows
Example:
const rows5to10 = df.slice(5, 10); // Rows 5-9
const first100 = df.slice(0, 100); // Rows 0-99Detect missing values.
Parameters:
columnName: string- Column to check
Returns: New DataFrame with boolean column {columnName}_isna
Example:
const result = df.isna('age');
result.show();
// name age age_isna
// 0 Alice 30 false
// 1 Bob - true
// 2 Charlie 35 falseDetect non-missing values.
Parameters:
columnName: string- Column to check
Returns: New DataFrame with boolean column {columnName}_notna
Example:
const result = df.notna('age');
result.show();
// name age age_notna
// 0 Alice 30 true
// 1 Bob - false
// 2 Charlie 35 trueDrop rows with missing values.
Parameters:
columnName: string- Column to check
Returns: New DataFrame without rows having null in specified column
Example:
// Remove rows where age is missing
const cleaned = df.dropna('age');
// Chain multiple dropna calls
const fullyClean = df
.dropna('age')
.dropna('city')
.dropna('score');Fill missing values.
Parameters:
columnName: string- Column to fillfillValue: any- Value to use for missing data
Returns: New DataFrame with filled values
Example:
// Fill missing ages with 0
const df2 = df.fillna('age', 0);
// Fill missing scores with mean
const meanScore = df.mean('score');
const df3 = df.fillna('score', meanScore);
// Fill missing strings
const df4 = df.fillna('city', 'Unknown');All string operations create a new DataFrame with the transformed column.
Convert strings to lowercase.
Example:
const df2 = df.strLower('email');
// alice@example.com → alice@example.com
// BOB@EXAMPLE.COM → bob@example.comConvert strings to uppercase.
Example:
const df2 = df.strUpper('product');
// widget → WIDGET
// gadget → GADGETRemove leading/trailing whitespace.
Example:
const df2 = df.strTrim('name');
// " Alice " → "Alice"
// "Bob" → "Bob"Check if string contains substring.
Parameters:
columnName: string- Column namepattern: string- Substring to search for
Returns: New DataFrame with boolean column {columnName}_contains
Example:
const result = df.strContains('email', '@example.com');
result.show();
// email email_contains
// 0 alice@example.com true
// 1 bob@company.com falseReplace substring.
Parameters:
columnName: string- Column nameold: string- Substring to replacenew: string- Replacement substring
Example:
const df2 = df.strReplace('product', 'Widget', 'Component');
// "Widget A" → "Component A"
// "Gadget B" → "Gadget B"Extract substring.
Parameters:
columnName: string- Column namestart: number- Start indexend: number- End index (exclusive)
Example:
const df2 = df.strSlice('name', 0, 3);
// "Alice" → "Ali"
// "Bob" → "Bob"Check if string starts with prefix.
Parameters:
columnName: string- Column nameprefix: string- Prefix to check
Returns: New DataFrame with boolean column {columnName}_startswith
Example:
const result = df.strStartsWith('product', 'Widget');
// product product_startswith
// 0 Widget A true
// 1 Gadget B falseCheck if string ends with suffix.
Parameters:
columnName: string- Column namesuffix: string- Suffix to check
Returns: New DataFrame with boolean column {columnName}_endswith
Example:
const result = df.strEndsWith('email', '.com');
// email email_endswith
// 0 alice@example.com true
// 1 bob@example.org falseGet string length.
Parameters:
columnName: string- Column name
Returns: New DataFrame with integer column {columnName}_len
Example:
const result = df.strLen('name');
// name name_len
// 0 Alice 5
// 1 Bob 3Absolute value.
Example:
const df2 = df.abs('temperature');
// -5 → 5, 10 → 10Round to N decimal places.
Parameters:
decimals?: number- Decimal places (default: 0)
Example:
const df2 = df.round('score', 1);
// 95.567 → 95.6Sum of column values.
Example:
const total = df.sum('sales');
console.log(total); // 15000Mean (average) of column values.
Example:
const avgAge = df.mean('age');
console.log(avgAge); // 28.5Minimum value.
Example:
const minScore = df.min('score');
console.log(minScore); // 72.3Maximum value.
Example:
const maxScore = df.max('score');
console.log(maxScore); // 98.5Standard deviation.
Example:
const stdDev = df.std('age');
console.log(stdDev); // 5.2Variance.
Example:
const variance = df.variance('score');
console.log(variance); // 27.04Median value (50th percentile).
Example:
const medianAge = df.median('age');
console.log(medianAge); // 30Quantile (percentile).
Parameters:
columnName: string- Column nameq: number- Quantile (0.0 to 1.0)
Example:
const q25 = df.quantile('score', 0.25); // 25th percentile
const q50 = df.quantile('score', 0.50); // 50th percentile (median)
const q75 = df.quantile('score', 0.75); // 75th percentile
const q90 = df.quantile('score', 0.90); // 90th percentileFrequency distribution.
Returns: DataFrame with columns: {columnName}, count
Example:
const counts = df.valueCounts('grade');
counts.show();
// grade count
// 0 A 5
// 1 B 3
// 2 C 2Correlation matrix.
Parameters:
columnNames: string[]- Columns to correlate
Returns: Correlation matrix DataFrame
Example:
const corr = df.corrMatrix(['math', 'science', 'english']);
corr.show();
// math science english
// math 1.00 0.85 0.72
// science 0.85 1.00 0.68
// english 0.72 0.68 1.00Rank values.
Parameters:
columnName: string- Column to rankmethod: 'average' | 'min' | 'max' | 'dense' | 'ordinal'- Ranking method
Returns: New DataFrame with column {columnName}_rank
Example:
const ranked = df.rank('score', 'average');
ranked.show();
// name score score_rank
// 0 Alice 95 1.0
// 1 Bob 87 3.0
// 2 Charlie 95 1.0 (tied, average)
// 3 Diana 82 4.0Rolling sum.
Parameters:
columnName: string- Column namewindowSize: number- Window size
Returns: New DataFrame with column {columnName}_rolling_sum
Example:
const df2 = df.rollingSum('sales', 3);
// [10, 20, 30, 40] → [null, null, 60, 90]Rolling mean (moving average).
Example:
const sma5 = df.rollingMean('price', 5); // 5-day SMA
const sma10 = df.rollingMean('price', 10); // 10-day SMARolling minimum.
Example:
const df2 = df.rollingMin('price', 3);Rolling maximum.
Example:
const df2 = df.rollingMax('price', 3);Rolling standard deviation (volatility).
Example:
const volatility = df.rollingStd('price', 20);Cumulative sum.
Example:
const cumSum = df.expandingSum('sales');
// [10, 20, 30] → [10, 30, 60]Cumulative mean.
Example:
const cumMean = df.expandingMean('score');
// [90, 80, 85] → [90, 85, 85]Sort by columns.
Parameters:
columnNames: string | string[]- Column(s) to sort byascending?: boolean | boolean[]- Sort order (default: true)
Returns: Sorted DataFrame
Example:
// Sort by single column
const df2 = df.sortBy('age'); // ascending
const df3 = df.sortBy('age', false); // descending
// Sort by multiple columns
const df4 = df.sortBy(['city', 'age']); // both ascending
// Mixed sort order
const df5 = df.sortBy(['city', 'age'], [true, false]);
// city ascending, age descendingJoin DataFrames (inner join).
Parameters:
other: DataFrame- DataFrame to join withon: string- Column name to join onhow?: 'inner'- Join type (default: 'inner')
Returns: Joined DataFrame
Example:
const customers = DataFrame.fromCSV(`id,name
1,Alice
2,Bob`);
const orders = DataFrame.fromCSV(`id,customer_id,total
101,1,100
102,2,200
103,1,150`);
const joined = orders.join(customers, 'id');Left outer join.
Example:
const result = df.leftJoin(other, 'id');
// Keeps all rows from df, matching rows from otherRight outer join.
Example:
const result = df.rightJoin(other, 'id');
// Keeps all rows from other, matching rows from dfFull outer join.
Example:
const result = df.outerJoin(other, 'id');
// Keeps all rows from both DataFramesCross join (Cartesian product).
Example:
const result = df.crossJoin(other);
// Every row from df × every row from otherGroup by column.
Returns: GroupedDataFrame for aggregation
Example:
const grouped = df.groupBy('department');
// Aggregate
const result = grouped.agg({
salary: 'mean',
age: 'mean'
});
result.show();
// department salary_mean age_mean
// 0 Engineering 95000 32.5
// 1 Marketing 82000 29.0Pivot table (long to wide).
Parameters:
index: string- Row index columncolumns: string- Column to pivotvalues: string- Values to aggregateaggFunc: 'sum' | 'mean' | 'min' | 'max' | 'count'- Aggregation function
Example:
const df = DataFrame.fromCSV(`store,product,sales
A,Widget,100
A,Gadget,80
B,Widget,110
B,Gadget,85`);
const pivoted = df.pivot('store', 'product', 'sales', 'sum');
pivoted.show();
// store Widget Gadget
// 0 A 100 80
// 1 B 110 85Unpivot table (wide to long).
Parameters:
idVars: string[]- Columns to keep as identifiersvalueVars: string[]- Columns to unpivotvarName: string- Name for variable columnvalueName: string- Name for value column
Example:
const df = DataFrame.fromCSV(`student,math,science
Alice,95,92
Bob,78,85`);
const melted = df.melt(['student'], ['math', 'science'], 'subject', 'score');
melted.show();
// student subject score
// 0 Alice math 95
// 1 Alice science 92
// 2 Bob math 78
// 3 Bob science 85Swap rows and columns.
Example:
const df2 = df.transpose();
// Rows become columns, columns become rowsStack columns into rows.
Example:
const stacked = df.stack();Unstack rows into columns.
Example:
const unstacked = df.unstack();Export DataFrame schema to Arrow format.
Returns: Arrow schema object (JSON)
Example:
const arrowSchema = df.toArrow();
console.log(arrowSchema);
// {
// schema: {
// fields: [
// { name: 'name', type: { name: 'utf8' }, nullable: true },
// { name: 'age', type: { name: 'int' }, nullable: false }
// ]
// }
// }Import DataFrame from Arrow schema.
Parameters:
arrowSchema: ArrowSchema- Arrow schema object
Returns: DataFrame
Example:
const schema = {
schema: {
fields: [
{ name: 'id', type: { name: 'int' }, nullable: false },
{ name: 'value', type: { name: 'floatingpoint' }, nullable: false }
]
}
};
const df = DataFrame.fromArrow(schema);Create lazy DataFrame for query optimization.
Returns: LazyDataFrame
Example:
const lazyDf = df.lazy();Add column selection to query plan.
Parameters:
columnNames: string[]- Columns to select
Returns: LazyDataFrame
Example:
const lazy = df.lazy().select(['name', 'age']);Add row limit to query plan.
Parameters:
n: number- Number of rows
Returns: LazyDataFrame
Example:
const lazy = df.lazy().limit(100);Execute optimized query plan.
Returns: DataFrame with results
Example:
const result = df.lazy()
.select(['name', 'age'])
.limit(10)
.collect(); // Execute now!
result.show();Free DataFrame memory (C ABI required).
Important: Always call free() when done with a DataFrame to prevent memory leaks.
Example:
const df = DataFrame.fromCSV(csv);
// Use DataFrame
df.show();
// Free memory
df.free();Pattern: Use try/finally for cleanup
let df;
try {
df = DataFrame.fromCSV(csv);
// Use DataFrame
const result = df.filter(row => row.get('age') > 30);
result.show();
// Free intermediate results
result.free();
} finally {
// Always free, even if error
if (df) df.free();
}All operations have full TypeScript definitions with JSDoc examples.
import { Rozes, DataFrame, RowRef } from 'rozes';
const rozes = await Rozes.init();
const DataFrame = rozes.DataFrame;
const df: DataFrame = DataFrame.fromCSV(csvString);
// Filter with type safety
const filtered: DataFrame = df.filter((row: RowRef) => {
const age = row.get('age');
return typeof age === 'number' && age > 30;
});
// Cleanup
filtered.free();
df.free();-
Lazy evaluation: Use for chained operations on large datasets
// Eager (slower) const result = df.select(['a', 'b']).head(10); // Lazy (faster) const result = df.lazy().select(['a', 'b']).limit(10).collect();
-
Projection pushdown: Select columns early
// Bad: Load all columns, then select const result = df.filter(pred).select(['a', 'b']); // Good: Select first (fewer columns to filter) const result = df.select(['a', 'b']).filter(pred);
-
Batch operations: Process in chunks for large datasets
const chunkSize = 10000; for (let i = 0; i < df.shape.rows; i += chunkSize) { const chunk = df.slice(i, i + chunkSize); // Process chunk chunk.free(); }
-
Memory management: Free intermediate results
const df1 = df.filter(pred1); const df2 = df1.filter(pred2); df1.free(); // Free intermediate result // Use df2 df2.free();
try {
const df = DataFrame.fromCSV(csvString);
// Operations that might fail
const result = df.filter(row => {
const age = row.get('age');
if (typeof age !== 'number') {
throw new Error('Invalid age type');
}
return age > 30;
});
result.show();
result.free();
df.free();
} catch (err) {
console.error('DataFrame error:', err.message);
}- README.md - Quick start guide
- examples/js/ - API showcase examples
- examples/nodejs/ - Real-world examples
- CHANGELOG.md - Version history
Last Updated: 2025-11-08 | Version: 1.3.0