They look very comparable. When I ran the algorithm and giving it 1000 starts under the red triangle my results were lightly different.
I wouldn't lose sleep over either choice.
3802_Screen Shot 2013-06-27 at 11.52.38 AM.png
3803_Screen Shot 2013-06-27 at 11.52.57 AM.png
How about a compromise with 2 replicates and 18 runs total? I have attached all three scenarios with the same simulated response model for you to see. Just run the Fit model scripts.
As you can see all approaches are comparable.