# Early Stopping

**URL:** <https://community.wandb.ai/t/early-stopping/422>\
**Category:** W&B Help\
**Tags:** wandb\
**Created:** [September 3, 2021, 9:39pm UTC](https://community.wandb.ai/t/early-stopping/422 "2021-09-03T21:39:56Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![max\_wasserman](https://avatars.discourse-cdn.com/v4/letter/m/7bcc69/32.png) [@max\_wasserman](https://community.wandb.ai/u/max_wasserman)\
**Post date:** [September 3, 2021, 9:39pm UTC](https://community.wandb.ai/t/early-stopping/422/1 "2021-09-03T21:39:56Z")

</div>

Hello All,

I’m configuring a hyper parameter sweep. I have training, validation, and test set.

I’d like to use the test\_loss as the final metric to optimize and val\_loss for early stopping.

I don’t see a place to specify a metric for early stopping. Does it default to the same metric specified for overall optimization (of hyper parameters)? If so, how can I change this?

Thanks!

---

<div class="post-metadata">

**Author:** ![\_scott](https://sea2.discourse-cdn.com/flex020/user_avatar/community.wandb.ai/_scott/32/95_2.png) [@\_scott](https://community.wandb.ai/u/_scott)\
**Post date:** [September 4, 2021, 10:00am UTC](https://community.wandb.ai/t/early-stopping/422/2 "2021-09-04T10:00:58Z")

</div>

It isn’t possible to have a different metric for hyperband early stopping and search strategy. [https://github.com/wandb/sweeps/blob/master/hyperband\_stopping.py#L176](https://github.com/wandb/sweeps/blob/master/hyperband_stopping.py#L176)

One workaround would be to use a search strategy that doesn’t require a metric like `random` or `grid` and then use `val_loss` as your metric for early stopping. You can then easily reconfigure the resulting parameter importance and parallel coordinate plots to show `test_loss` in your dashboard.

If you would would like this feature, you can file a feature request on our client repo issues.

> **[GitHub - wandb/client: 🔥 A tool for visualizing and tracking your machine...](https://github.com/wandb/client)**
>
> 🔥 A tool for visualizing and tracking your machine learning experiments. This repo contains the CLI and Python API. - GitHub - wandb/client: 🔥 A tool for visualizing and tracking your machine learn...

---

<div class="post-metadata">

**Author:** ![charlesfrye](https://sea2.discourse-cdn.com/flex020/user_avatar/community.wandb.ai/charlesfrye/32/52_2.png) [@charlesfrye](https://community.wandb.ai/u/charlesfrye)\
**Post date:** [September 5, 2021, 1:21am UTC](https://community.wandb.ai/t/early-stopping/422/3 "2021-09-05T01:21:46Z")

</div>

Welcome to the forum, @max_wasserman! Great first question.

While I can see why it might be good for us to add the ability to separate the early-stopping metric from the Bayesian optimization metric, **I would strongly caution against using the test loss in any step of the process** – whether its the optimization of parameters (obviously a no-no!) or the optimization of hyperparameters. The PyTorch Lightning docs even say that you should only call `.test` [“[o]nly right before publishing your paper or pushing to production”](https://pytorch-lightning.readthedocs.io/en/latest/common/trainer.html#testing).

The purpose of metrics measured on the test set is to reflect, as veridically as possible, the performance of the model on more data drawn from the same distribution, which we in turn hope reflects the performance of the model on data in production. Selecting hyperparameters based on the test set breaks the “information wall” (more technically, the conditional independence relation) between the test data and the model’s parameters that make the test set useful for getting unbiased estimates of true generalization performance.

There is at least some indication that the use of fixed validation and test sets has led the ML field as a whole to “overfit”, in the sense of over-estimation of true generalization performance:

> **[Do CIFAR-10 Classifiers Generalize to CIFAR-10?](https://arxiv.org/abs/1806.00451)**
>
> Machine learning is currently dominated by largely experimental work focused
> on improvements in a few key tasks. However, the impressive accuracy numbers of
> the best performing models are questionable because the same test sets have
> been used to...

> **[Do ImageNet Classifiers Generalize to ImageNet?](http://proceedings.mlr.press/v97/recht19a.html)**
>
> We build new test sets for the CIFAR-10 and ImageNet datasets. Both benchmarks have been the focus of intense research for almost a decade, raising the danger of overfitting to excessively re-used ...

---

<div class="post-metadata">

**Author:** ![max\_wasserman](https://avatars.discourse-cdn.com/v4/letter/m/7bcc69/32.png) [@max\_wasserman](https://community.wandb.ai/u/max_wasserman)\
**Post date:** [September 5, 2021, 7:17pm UTC](https://community.wandb.ai/t/early-stopping/422/4 "2021-09-05T19:17:08Z")

</div>

> [@charlesfrye](#):
>
> electing hyperparameters based on the test set breaks the “information wall” (more technically, the conditional independence relation) betw

Thanks so much for the responses.

In this case I am using synthetic data (I can generate a lot it cheaply). I used the names val/test\_loss instead of validation set 1 (for early stopping) and validation set 2 (for bayes optimization) for simplicity. I will generate more data after this (my true test set) for final unbiased estimation of generalization.

It seems the best solution at the moment is to simply do what @_scott recommended: use a random search and log ‘test\_loss’ (actually validation set 2) for viz later.

PS is this the preferred location/forum where I should post technical questions of this kind? The GitHub page refers to a slack group that appears to be closed.

---

<div class="post-metadata">

**Author:** ![bhutanisanyam1](https://sea2.discourse-cdn.com/flex020/user_avatar/community.wandb.ai/bhutanisanyam1/32/18_2.png) [@bhutanisanyam1](https://community.wandb.ai/u/bhutanisanyam1)\
**Post date:** [September 5, 2021, 7:34pm UTC](https://community.wandb.ai/t/early-stopping/422/5 "2021-09-05T19:34:38Z")

</div>

> [@max\_wasserman](#):
>
> PS is this the preferred location/forum where I should post technical questions of this kind? The GitHub page refers to a slack group that appears to be closed.

Yes, this is the place, Thanks for checking!

We really want to make sure our community enjoys the forums so we’re silently moving from slack to discourse and we’ll be making the announcement soon once we’re confident the forums are all setup 🙂

---

<div class="post-metadata">

**Author:** ![charlesfrye](https://sea2.discourse-cdn.com/flex020/user_avatar/community.wandb.ai/charlesfrye/32/52_2.png) [@charlesfrye](https://community.wandb.ai/u/charlesfrye)\
**Post date:** [September 5, 2021, 7:54pm UTC](https://community.wandb.ai/t/early-stopping/422/6 "2021-09-05T19:54:39Z")

</div>

> [@max\_wasserman](#):
>
> I will generate more data after this (my true test set) for final unbiased estimation of generalization.

Ah okay, if you’ve got an actual unbiased test set, then you’re golden. I’d be interested to hear more about your project!

And yes, as @_scott points out, if you aren’t using Bayesian optimization, the choice of `metric` won’t impact the behavior of your search. `random` is actually a pretty good choice for HPO, competitive with `bayes` in my and others’ experience – and less prone to error/misconfiguration. Also BTW, [the `early_terminate` feature uses HyperBand](https://docs.wandb.ai/guides/sweeps/configuration#early_terminate), which is more aggressive than the usual early stopping folks learn about in an ML class, based on stopping training when you see increasing validation set error. That style of early stopping is best delegated to the ML framework you’re using.

Thanks for pointing out the issue with the Slack link. As @bhutanisanyam1 said, we are moving discussion to this forum, but that link should’ve still been in operation anyway. Will fix it shortly.

---

<div class="post-metadata">

**Author:** ![max\_wasserman](https://avatars.discourse-cdn.com/v4/letter/m/7bcc69/32.png) [@max\_wasserman](https://community.wandb.ai/u/max_wasserman)\
**Post date:** [September 5, 2021, 11:48pm UTC](https://community.wandb.ai/t/early-stopping/422/7 "2021-09-05T23:48:41Z")

</div>

I’m doing some graph learning work (inputs are graphs, labels are graphs). Submitting paper soon, so I’ll post it to one of these forums after!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex020/uploads/wandb/original/1X/366b649231631dbab896843020da0056074ac79d.png) [@system](https://community.wandb.ai/u/system)\
**Post date:** [April 20, 2022, 6:02pm UTC](https://community.wandb.ai/t/early-stopping/422/8 "2022-04-20T18:02:07Z")

</div>

This topic was automatically closed 60 days after the last reply. New replies are no longer allowed.
