Xiaomi TabLDM: Open-source data model tops benchmark

Xiaomi has released an open-source large foundational model for structured tabular data called Xiaomi-TabLDM. The model is designed to handle both classification and regression tasks across various tabular datasets without needing task-specific retraining or tuning. It aims to simplify deployment in industries where tabular data is common, such as finance, healthcare, manufacturing, and logistics.

Model Design and Pretraining Approach

Xiaomi-TabLDM is pretrained solely on synthetic datasets generated by structural causal models (SCM). These datasets cover various scales, variable types, dependencies, and functional relations to broaden the model’s training distribution. The architecture includes dual-stream feature grouping for finer relational modeling, lightweight attention residuals, and a sparse mixture-of-experts mechanism to expand capacity without large computational overhead.

For inference, the model employs Test-Time Scaling. This technique improves prediction accuracy by increasing computing effort during inference while keeping pretrained parameters fixed.

Benchmark Performance and Efficiency

The model ranks in the top tier across four major public benchmarks. It secured first place in regression on the OpenML-CTR23 leaderboard, and second place on TALENT, TabArena, and BCCO regression tasks. For classification, Xiaomi-TabLDM topped the TALENT binary classification task.

In the TabArena regression benchmark, Xiaomi-TabLDM balanced performance and speed effectively. It achieved the second highest Elo rating while reducing training time by 82% and inference time by 68% compared to the top model TabFM. (Elo ratings here measure relative performance ranking.)

Industrial Validation and Practical Benefits

Tests in real-world industrial scenarios demonstrated Xiaomi-TabLDM’s superior accuracy and adaptability over traditional machine learning models like XGBoost. In material property prediction, the model improved prediction accuracy by 130% without any fine-tuning, cutting invalid sample tests by about 90%. For part weight predictions, it reduced error sample rates by 31%. In production component prediction, mean error dropped 54%, with a further 62% error reduction after introducing only about 30 new samples to adapt to new production conditions.

Availability and Access

Xiaomi-TabLDM is compatible with scikit-learn and can be installed via pip. Both the model’s code and pretrained weights are publicly available:

This release marks Xiaomi’s effort to establish a foundational model approach for structured tabular data, reducing the burden of retraining models for distinct datasets and improving prediction efficiency and accuracy across multiple domains.

Users interested in trying Xiaomi-TabLDM can access the latest resources and pretrained models online. Feature and update availability may vary depending on integration and use case complexity.

Source: ithome.com

MemeOS Enhancer Download
Avatar for Emir Bardakçı

Emir Bardakçı

Co-founder & HyperOS Expert

Keeping a pulse on Xiaomi, HyperOS, and the Android world. Tech enthusiast, photography lover, and detailed reviewer.

Leave a Reply

Your email address will not be published. Required fields are marked *