It is very, very hard to enforce behavior to an optimization system just with rewards / penalties and no explicit constraints. Which is why in manufacturing we use MPC, not RL (or use them within a system that can outright reject their recommendations if dangerous).
There will be always cases that sacrificing one direction (operating rules) can improve the other one (profit).
There will be always cases that sacrificing one direction (operating rules) can improve the other one (profit).