Fair expectimax

Look a few moves ahead, take the best column on your own turns, and average over every sampled disc the game might deal.

Scores visible board traits, adds them together, and plays the highest-valued column.

Lifetime objective

Judge a move by how much longer the game will still last, rather than by how many points it scores right now.

Value and policy learning

Instead of searching ahead, train a model on past games to judge a board or pick a column, and learn why that kept failing.