Skip to content

AutoML experiments in non declarative style not working #6446

Description

@thoron

System Information (please complete the following information):

  • OS & Version: Windows 11
  • ML.NET Version: 2.0
  • .NET Version: .NET6

Describe the bug
Running the "old" AutoML experiments does not work for all trainers. Using CreateMulticlassClassificationExperiment or CreateRegressionExperiment instead of the new declarative style will result in an exception (possibly due to custom schemas).

An exception will be raised in the new SweepablePipeline:

System.NullReferenceException: Object reference not set to an instance of an object.
   at Microsoft.ML.AutoML.SweepablePipeline..ctor(Dictionary`2 estimators, Entity schema, String currentSchema)
   at Microsoft.ML.AutoML.SweepablePipeline.AppendEntity(Boolean allowSkip, Entity entity)
   at Microsoft.ML.AutoML.AutoCatalog.MultiClassification(String labelColumnName, String featureColumnName, String exampleWeightColumnName, Boolean useFastForest, Boolean useLgbm, Boolean useFastTree, Boolean useLbfgs, Boolean useSdca, FastTreeOption fastTreeOption, LgbmOption lgbmOption, FastForestOption fastForestOption, LbfgsOption lbfgsOption, SdcaOption sdcaOption, SearchSpace`1 fastTreeSearchSpace, SearchSpace`1 lgbmSearchSpace, SearchSpace`1 fastForestSearchSpace, SearchSpace`1 lbfgsSearchSpace, SearchSpace`1 sdcaSearchSpace)
   at Microsoft.ML.AutoML.MulticlassClassificationExperiment.CreateMulticlassClassificationPipeline(IDataView trainData, ColumnInformation columnInformation, IEstimator`1 preFeaturizer)
   at Microsoft.ML.AutoML.MulticlassClassificationExperiment.Execute(IDataView trainData, ColumnInformation columnInformation, IEstimator`1 preFeaturizer, IProgress`1 progressHandler)
   at Microsoft.ML.AutoML.MulticlassClassificationExperiment.Execute(IDataView trainData, String labelColumnName, String samplingKeyColumn, IEstimator`1 preFeaturizer, IProgress`1 progressHandler)

To Reproduce

var experimentSettings = new MulticlassExperimentSettings
{
  MaxExperimentTimeInSeconds = 30,
  OptimizingMetric = MulticlassClassificationMetric.MacroAccuracy
};
experimentSettings.Trainers.Clear();
experimentSettings.Trainers.Add(MulticlassClassificationTrainer.LbfgsMaximumEntropy);
ctx.Auto().CreateMulticlassClassificationExperiment(experimentSettings).Execute(trainDv);

Where the schema is of a custom type:

var schemaDef = SchemaDefinition.Create(typeof(ModelInput));
schemaDef["Features"].ColumnType = new VectorDataViewType(NumberDataViewType.Single, numberOfFeatures);
schemaDef.Remove(schemaDef["LabelFeaturized"]); // removed when not applicable
public class ModelInput
{
  public uint Label;
  public float LabelFeaturized;
  public float[] Features;
}

Expected behavior
No regression expected for CreateMulticlassClassificationExperiment and CreateRegressionExperiment.

Activity

  1. ghost added
    untriagedNew issue has not been triaged
    on Nov 10, 2022
  2. changed the title [-]AutoML experiments not working[/-] [+]AutoML experiments in non declarative style not working[/+] on Nov 10, 2022
  3. LittleLittleCloud commented on Nov 10, 2022

    @LittleLittleCloud
    Member

    Thanks for reporting this bug. It's because LbfgsMaximumEntropy is not used when constructing sweepable pipeline for multiclass classification (

    private SweepablePipeline CreateMulticlassClassificationPipeline(IDataView trainData, ColumnInformation columnInformation, IEstimator<ITransformer> preFeaturizer = null)
    )

    We'll push a fix for this bug, in the meanwhile, you can add other trainer (like lightGbm) as a work-around

  4. thoron commented on Nov 11, 2022

    @thoron
    ContributorAuthor

    Thanks for reporting this bug. It's because LbfgsMaximumEntropy is not used when constructing sweepable pipeline for multiclass classification (

    private SweepablePipeline CreateMulticlassClassificationPipeline(IDataView trainData, ColumnInformation columnInformation, IEstimator<ITransformer> preFeaturizer = null)

    )

    We'll push a fix for this bug, in the meanwhile, you can add other trainer (like lightGbm) as a work-around

    Thank you for your answer! We will downgrade until the fix is deployed.

  5. added this to the ML.NET 3.0 milestone on Nov 28, 2022
  6. ghost removed
    untriagedNew issue has not been triaged
    on Nov 28, 2022
  7. zeroskyx commented on Dec 3, 2022

    @zeroskyx

    Any chance of getting this fix rolled out before ML.NET 3.0 as an update to ML.NET 2.0? Projects depending on this would otherwise need to stay at 0.7.1/ 0.19.1 :(

  8. michaelgsharp commented on Dec 5, 2022

    @michaelgsharp
    Contributor

    @luisquintanilla how does the suggestion by @zeroskyx sound for a servicing release?

  9. thoron commented on Dec 6, 2022

    @thoron
    ContributorAuthor

    I believe it should be in a servicing release as it will produce a runtime exception when simply upgrading and using a previously functional workflow (without any syntax or compilation warnings / errors).

  10. zeroskyx commented on Dec 23, 2022

    @zeroskyx

    Possibly related to #6529

  11. added a commit that references this issue on Jan 5, 2023
    bc250df
  12. ghost removed on Jan 5, 2023
  13. ghost locked as resolved and limited conversation to collaborators on Feb 5, 2023
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions