Skip to content
All projects

Nifty 200 Portfolio Optimisation

FinanceMachine Learning

Built with

  • python
  • pandas
  • yfinance
  • scikit-learn

Overview

Grouping Nifty 200 stocks by their fundamentals, then optimising each group on its own, beat a market-cap-weighted Nifty 200 on every risk-adjusted measure we tested. Index funds weight stocks by size, so a few large companies dominate whatever their risk. We asked whether letting the data find stocks that resemble each other, and weighting them by risk and return instead, would do better. Over a 3-year window all 20 clusters beat the benchmark, with an average Sharpe ratio of 1.47 against -0.67. The project was built by a team of five in Python (pandas, scikit-learn, SciPy).

Data and preparation

We took the 200 stocks of the Nifty 200 from the NSE and pulled prices and financials from Yahoo Finance. Each stock is described by 7 fundamentals covering profitability (EBIT, ROE, ROA), valuation (P/E), leverage (debt-to-equity), size (market cap) and market risk (beta). Missing metrics were filled with the median, which is robust to outliers, and missing prices with the previous price. All metrics were then Z-score standardised, so that market cap in crores does not drown out a beta near 1 when K-Means measures distance.

Methodology

K-Means clustering grouped the standardised stocks into 20 clusters, with K chosen by the elbow method over K = 1 to 40 (inertia flattened out beyond 20). Inside each cluster we found the weights that maximise the Sharpe ratio, using Markowitz optimisation (long-only, weights summing to 1) on annualised returns and the covariance of daily returns. Each cluster was then scored on the Sharpe, Treynor and Reward-to-Risk ratios and on 95% Value-at-Risk, and compared with a Nifty 200 benchmark weighted by market cap.

Choosing the timeframe

To avoid fitting one market mood, we repeated the analysis over 3 months, 6 months, 1 year and 3 years and averaged each ratio across the 20 clusters. The 3-year window won on every metric: a Sharpe of 1.47 (against 0.90 for one year and roughly zero for the shorter windows), a Treynor of 0.29, a Reward-to-Risk of 1.47, and the fewest zero-weight stocks (23), so the optimiser could use more of the universe. Short windows catch abnormal events, while three years spans growth, correction and volatile phases.

Results

Over the 3 years, the market-cap-weighted benchmark scored a Sharpe of -0.67, a Treynor of -0.11 and a Reward-to-Risk of -0.32. The 20 optimised clusters averaged 1.47, 0.29 and 1.47, and every cluster beat the benchmark. The best was Cluster 7, 26 large, low-beta (0.67) stocks with a Sharpe of 2.69, where three holdings (Lupin, Dixon Technologies and BSE) carry about two-thirds of the weight. Risk varied widely: 95% VaR averaged 0.42, from 0.22 for Cluster 1 (mature, stable large caps) to 0.70 for Cluster 13 (highly leveraged firms), while loss-making Cluster 2 was the weakest.

What it means for investors

Building portfolios from fundamentals rather than size gave better risk-adjusted outcomes, because each cluster holds stocks that behave alike and the optimiser spreads weight efficiently inside it. The clusters also read as a menu by risk appetite: Clusters 0 and 1 (mature, low-beta, low-leverage large caps) suit conservative investors, Cluster 7 and its peers offer the strongest risk-adjusted returns, and distressed groups such as Clusters 2 and 18 are best avoided. It is a framework for screening and allocation, not investment advice.

Limitations

Three judgement calls shape the results. Median imputation filled metrics missing for a few stocks, which is robust but still an assumption. The choice of 20 clusters came from reading the elbow curve, which is subjective, and another analyst might pick a different K. The 3-year window smooths out noise but may reflect today's fast-moving market less well than a shorter one. Weights were also optimised and scored on the same period (in-sample), so the natural next step is testing them on unseen data.