[Roadmap] Multiple outputs.
Updates
In v3.4.0, the hist tree method is considered feature complete for the vector leaf.
Context
Since XGBoost 1.6, we have been working on having multi-output support for the tree model. In 2.0, we implemented the initial version of the vector-leaf-based multi-output model. This issue serves as a tracker for future development and related discussion. The original feature request is here: https://github.com/dmlc/xgboost/issues/2087 . The related features are for vector-leaf, not for general multi-output.
Feel free to share your suggestions or make related feature requests in the comments.
Implementation Optimization
- Use f-order for the gradient. Currently, the gradient has one column for each target but is written in C-order. The transformation takes about one-fifth of the training time. (#9508)
- Use f-order for the custom objective. (#9089)
- Improve array type dispatching by moving the dispatch logic from per-element to per-array. This enables us to have a more efficient custom objective interface. (#9090)
Algorithmic Optimization
We are still looking for potential algorithmic optimization for vector-leaf and here's the pool of candidates. We need to survey all available options. Feel free to share if you have ideas or paper recommendations.
- Sketch boost. (#11798, #11922)
- https://arxiv.org/abs/2201.06239 (#11798, #11922)
- Extra tree.
(#11798)
GPU Implementation
- Evaluation (#11781, #11883)
- Histogram (#11781, #11855)
- Prediction (#11752)
- Prediction cache. (#11862)
- Model (#11277)
- Partition. (#11789)
- Gradient sampling.
Documentation
- Derive the approximated Hessian in the context of boosting trees.
Multi-task
- Multi-task xgboost. This is not yet decided. I think it's wise to at least do some exploration before forging the rest of the implementation since we will have a very different interface if we need to consider multi-task. Related: https://github.com/dmlc/xgboost/issues/7693 .
Features
- Tree SHAP
- Plotting (#10093)
- Model text dump (JSON, txt, graphviz) (#10093, #11747)
- Tree data frame. (#12293)
- Categorical feature. (#12072, #12276, #12299, #12305)
- Interaction constraints (#12294)
- Monotonic constraints (#12341)
- Subsample.
- Column sampling.
- Approx tree method
- Exact tree method
- Loss weight
- Feature importance (be careful with tree index) (#10700)
- Intercept. (#11656)
- dart (#12340)
Learning to rank
We can have a ranking model to consider multiple criteria. This might require multi-task to be supported.
Quantile regression
Distributed
- Dask (#12292)
- PySpark
- Spark
- Flink?
- Federated (https://github.com/dmlc/xgboost/pull/9171)
Binding
- R (https://github.com/dmlc/xgboost/pull/9526)
- Scala
- Python
- Java
- C
HPO
- Check compatibility with major HPO frameworks.
Other extensions
- Sparse label. (multi-label classification optimization)
- Missing label.
- Early stopping for each target?
Applications
Benchmarks
- Collection of datasets for future comparison.
Source: dmlc/xgboost