Browse
We’re here to help

Find guidance on Author Services

Search
Browse
We’re here to help

Find guidance on Author Services

Home
All Journals
Journal of Computational and Graphical Statistics
List of Issues
Volume 30, Issue 3
MIP-BOOST: Efficient and Effective L0 Fe ....

Search in:

Advanced search

Journal of Computational and Graphical Statistics Volume 30, 2021 - Issue 3

Submit an article Journal homepage

459

Views

CrossRef citations to date

Altmetric

Dimensionality Reduction, Regularization, and Variable Selection

MIP-BOOST: Efficient and Effective L₀ Feature Selection for Linear Regression

Ana Kenneya Department of Statistics, Penn State University, University Park, PA;; Correspondence[email protected]

https://orcid.org/0000-0002-2209-2431 View further author information

Francesca Chiaromontea Department of Statistics, Penn State University, University Park, PA;; ;b Institute of Economics & EMbeDS, Sant’Anna School of Advanced Studies, Pisa, Italy;; View further author information

Giovanni Felicic Istituto di Analisi dei Sistemi ed Informatica, Consiglio Nazionale delle Ricerche, Rome, ItalyView further author information

Pages 566-577 | Received 19 Sep 2019, Accepted 14 Oct 2020, Published online: 04 Jan 2021

Cite this article
https://doi.org/10.1080/10618600.2020.1845184
CrossMark

Sample our Computer Science journals, sign in here to start your access, latest two full volumes FREE to you for 14 days

Full Article
Figures & data
References
Supplemental
Citations
Metrics
Reprints & Permissions
Read this article /doi/full/10.1080/10618600.2020.1845184?needAccess=true

Abstract

Recent advances in mathematical programming have made mixed integer optimization a competitive alternative to popular regularization methods for selecting features in regression problems. The approach exhibits unquestionable foundational appeal and versatility, but also poses important challenges. Here, we propose MIP-BOOST, a revision of standard mixed integer programming feature selection that reduces the computational burden of tuning the critical sparsity bound parameter and improves performance in the presence of feature collinearity and of signals that vary in nature and strength. The final outcome is a more efficient and effective L₀ feature selection method for applications of realistic size and complexity, grounded on rigorous cross-validation tuning and exact optimization of the associated mixed integer program. Computational viability and improved performance in realistic scenarios is achieved through three independent but synergistic proposals. Supplementary materials including additional results, pseudocode, and computer code are available online.

Keywords:

Cross-validation
Feature selection
LASSO
Mixed integer optimization
Regression
Whitening

Supplementary Materials

Supplement-MIP-BOOST.pdf Contains pseudo code, additional simulation results, and a convergence proof sketch for bisection with feelers.

MIPBOOSTex.jl: Julia code running a simulated example of the methods described in the article. Contains functions for each component for general use.

Acknowledgments

We thank Matthew Reimherr for useful discussions and comments.

Notes

¹ Calculations were performed on the Roar Supercomputer at Penn State, using the basic memory option on the ACI-B cluster with an Intel Xeon 24 core processor at 2.2 GHz and 128 GB of RAM. The multi-thread option in Gurobi was limited to a maximum of 5 threads for consistency across settings.

² For SNR 10, MIP-BOOST was run with a more conservative threshold set for BF; a quick run of the LASSO selecting a very large number of features can inexpensively diagnose the need to implement a restrictive BF. However, we stress that BF used in MIP-BOOST does not suffer from the same lack of specificity that hinders LASSO at high SNRs. When the SNR increases the true k₀ becomes in fact easier to identify for MIP; it is simply that the BF search may be pushed toward the right of the k range if seemingly large differences in CVMSE are not “ignored” by setting a stricter threshold. In , the elbow at 10 is clearly denoted when $SNR = 10$ , but the strong signal leads to distinct drops in the CVMSE also at large k values. A more restrictive threshold easily improves solution quality. In contrast, using a more conservative tuning for LASSO (e.g., the parsimonious LSD), still fails to increase specificity.

Additional information

Funding

This work was partially funded by the NIH B2D2K training grant and the Huck Institutes of the Life Sciences of Penn State, and by NSF grant DMS-1407639. Computation was performed on the Roar Supercomputer at Penn State University.

Log in via your institution

Access through your institution

Log in to Taylor & Francis Online

Shibboleth

Log in to Taylor & Francis Online

Username Password

Forgot password?

Keep me logged in (not suitable for shared devices).

You will otherwise be logged out automatically, after a limited period, and will need to log in again.

Restore content access

Restore content access for purchases made as guest

Purchase options * Save for later Item saved, go to cart

PDF download + Online access

48 hours access to article PDF & online version
Article PDF can be downloaded
Article PDF can be printed

USD 61.00 Add to cart

PDF download + Online access - Online Checkout

Issue Purchase

30 days online access to complete issue
Article PDFs can be downloaded
Article PDFs can be printed

USD 180.00 Add to cart

Issue Purchase - Online Checkout

* Local tax will be added as applicable

Related Research

People also read lists articles that other readers of this article have read.

Recommended articles lists articles that we recommend and is powered by our AI driven recommendation engine.

Cited by lists all citing articles based on Crossref citations.
Articles with the Crossref icon will open in a new tab.

People also read
Recommended articles
Cited by

To cite this article:

Reference style: APA Chicago Harvard

Citation copied to clipboard

Reference styles above use APA (6th edition), Chicago (16th edition) & Harvard (10th edition)

Download citation

Download a citation file in RIS format that can be imported by citation management software including EndNote, ProCite, RefWorks and Reference Manager.

Choose format: RIS BibTex RefWorks Direct Export

Choose options: Citation Citation & abstract Citation & references

Information for

Authors
R&D professionals
Editors
Librarians
Societies

Open access

Overview
Open journals
Open Select
Dove Medical Press
F1000Research

Opportunities

Reprints and e-prints
Advertising solutions
Accelerated publication
Corporate access solutions

Help and information

Help and contact
Newsroom
All journals
Books

Keep up to date

Sign me up

Taylor and Francis Group Facebook page

Taylor and Francis Group X Twitter page

Taylor and Francis Group Linkedin page

Taylor and Francis Group Youtube page

Taylor and Francis Group Weibo page

Registered in England & Wales No. 3099067
5 Howick Place | London | SW1P 1WG

Your download is now in progress and you may close this window

Did you know that with a free Taylor & Francis Online account you can gain access to the following benefits?

Choose new content alerts to be informed about new research of interest to you
Easy remote access to your institution's subscriptions on any device, from any location
Save your searches and schedule alerts to send you new results
Export your search results into a .csv file to support your research

Have an account?
Login now Don't have an account?
Register for free

Login or register to access this feature

Have an account?
Login now Don't have an account?
Register for free

Choose new content alerts to be informed about new research of interest to you
Easy remote access to your institution's subscriptions on any device, from any location
Save your searches and schedule alerts to send you new results
Export your search results into a .csv file to support your research

MIP-BOOST: Efficient and Effective L0 Feature Selection for Linear Regression

Abstract

Supplementary Materials

Acknowledgments

Notes

Additional information

Funding

Log in via your institution

Log in to Taylor & Francis Online

Log in to Taylor & Francis Online

Restore content access

Related Research

To cite this article:

Download citation

Information for

Open access

Opportunities

Help and information

Keep up to date

Your download is now in progress and you may close this window

Login or register to access this feature

MIP-BOOST: Efficient and Effective L₀ Feature Selection for Linear Regression