I explain how I set the mixer for us to be happy with the monitors and ask him to try and respect that if he can.
This is essentially saying, thanks for coming and don't touch anything.
What you describe suggest the main mix probably improved as the monitor mix got worse. To mitigate this problem, the sound must be processed separately for the mains and monitors.
The signal flow in mixers varies. Monitors are usually run off Aux sends and should be set to pre fader. Pre fader allows the faders to be adjusted to make the main mix sound as good as possible, without changing the balance of instruments in the monitors.
Ideally you would also have different EQ settings for the house mix and monitors. Unfortunately channels only have one set of EQ controls. Better boards will allow the sends to be set pre or post EQ. With pre the sound going to the monitors is not EQ'ed (flat). With post, the channel EQ is global, meaning it applies to both the mains and monitors. There are different costs and benefits of each configuration.
Neither is a good solution when mic'ing low volume acoustic instruments because the monitors and mains are likely to have different problem frequencies that want to feedback.
Also the tone of the instruments can sound significantly different on stage VS out in the house. Ideally the sound of each instrument should be tailored separately for different listening locations.
Larger systems tend to have a versatile and powerful EQ on each output. These EQs are global to the mix, meaning they effect the entire mix. Typically 31 band graphic EQs or fully parametric EQ are used for this. These EQ can be used so the tonal balance of each mix is more alike (system EQ) and they can also be used to suppress feedback at different frequencies in each mix (utility EQ). What they can't do is adjust the tone of one instrument without impacting to tone of other instruments. The choice of mic and how you aim it can be helpful here.
Another way to manage the need for separate stage and house EQ is to dual assign the channels. For example if you have a 16 channel board and only use 8 channels, you can send mic 1 to channels 1 and 8, mic 2 to channels 2 and 9, mic 3 to channels 3 and 10, etc. Then you use channels 1 thru 8 for the house mix and channels 9 thru 16 for the monitors. This approach requires more resources and is more confusing and time intensive. Also, keep in mind the audio tech can't hear what you hear on stage. So there are some efficacy problems with applying EQ to the monitor mixes. One way around this would be for the audio tech to use a wireless tablet so he/she can come up on stage to dial in the sound of individual instruments to the monitors (during sound check).
If you are getting paid, I would argue that providing the best possible sound to the audience should be the top priority, but sound quality for the musicians is a close second. Unfortunately we have competing needs here, so some degree of compromise is usually required. This often translates to: As good as possible for the audience, and as bad as can be tolerated by the musicians. The key point is there is only so much sound quality to spread around, and often you can't improve the sound for one set of listeners without degrading the quality for the others. If the musicians insist on the best possible quality on stage, it tends to result in poor sound quality for the audience, and of course the audio tech is usually blamed.
Hopefully this gives you a better sense of some typical live-sound problems and work arounds. Ultimately live sound is more about achieving the best possible compromise, rather than pursuing perfection. It's rare for everyone to be happy with the results. Sounds like you hit the mark.