tsetick: A Python Library for Parsing and Querying Nikkei NEEDS Tick Data from the Tokyo Stock Exchange
Kazumi Li, Masataka Hayashi, Teruo Nakatsuma, Peter Romero
Abstract
Tick-level trade-and-quote data for the Tokyo Stock Exchange is distributed through the Nikkei NEEDS service as thousands of zipped CSV archives spanning four data types with era-dependent schemas and Japanese-language layouts. We present tsetick, an open-source Python library that converts these raw archives into clean, typed Polars DataFrames and a Hive-partitioned Parquet store queryable through DuckDB. The library offers two access paths sharing one parse-and-clean core: a one-shot reader that returns a ticker- and time-filtered DataFrame directly from raw ZIP files, and a two-stage ingest-then-query pipeline with resume-safe, memory-aware parallel ingestion, part-pruning, and a materialized intraday time key for row-group pruning. The engineering, more than the parsing, is what the library contributes: ingestion runs in per-date atomic units whose completion is recorded by coverage markers rather than inferred from file existence, writes stream in bounded morsels so that peak memory is independent of trading-day size (24.5 GB to 2.4 GB on the worst measured day), a RAM-aware process pool sizes itself to available memory, and part-pruning opens only the archive parts a ticker can occupy. Full English and Japanese column definitions ship for all four types, and a translation layer maps yfinance, Polygon, and ccxt names onto their tsetick equivalents. In benchmarks on a commodity 16-thread workstation, parsing a representative 4.8-million-row archive part, one of a trading day's nine parts, is 59.8x faster than the original pandas prototype (34.3x against an engine-matched pandas baseline), and a single-ticker time-window query from the store completes roughly 410x faster than a pandas scan of the equivalent CSV. tsetick is available on PyPI (pip install tse-tick) under the MIT license.
Create a lesson
Related papers
Price manipulation in nonlinear transient impact models: rigidity before memory and complete positivity after memory
Minhyeok Lee
Metaorder modelling and identification from public data
Ezra Goliath, Tim Gebbie
The Convergence Rate of Stochastic Tracking with Application to Optimal Execution
Marcel Nutz, Moritz Voss
Equilibrium in closed constant-function market maker economies
Muqiao Huang, Ruodu Wang, Yiyun Wang
Short-horizon mean reversion in cryptocurrency markets: a matched cross-market measurement
Nadav A. Kitron, Jonathan M. Wengrowicz
Concentrated Liquidity Provision: a Reinforcement Learning Perspective
Georgios Chionas, Charalampos Kleitsikas, Stefanos Leonardos et al.