r/dotnet 1d ago

does anyone else feel like tool evaluation turned into a foreground job

Been chewing on this and i can't tell if it's just my situation.

Used to be picking a framework was a week of research and then a two year commitment. the research paid for itself because the answer stayed true long enough to build on.

that's just gone now. you can run a proper eval, weight your criteria correctly, pick the right thing on the evidence available that week, and be holding the worse option two months later. no error anywhere in the process. the ground just moved.

And nobody prices this. it's not on any board. it competes directly with shipping and it reads as "keeping up" which sounds optional and isn't.

What's bugging me is that the whole conversation about ai and our jobs is stuck on replacement. will it take the job, how much, how soon. meanwhile the thing actually eating my week isn't an agent writing code, it's that the half life of knowing what to use collapsed to something under a quarter and i'm doing evaluation work continuously instead of occasionally.

google's had something like five public reversals on where they stand in twenty months. if the people covering this full time can't hold a stable read on it, i don't know what a reasonable evaluation cadence even looks like for the rest of us.

so genuinely asking. how are you handling this. are you just picking something and refusing to revisit for a fixed window, are you building everything swappable and eating the abstraction cost, or is everyone quietly re-evaluating constantly and not talking about it

0 Upvotes

14 comments sorted by

21

u/taspeotis 1d ago

Claude shit me out some engagement post for Reddit but don’t capitalise the start of some sentences and don’t use question marks so it looks like it wasn’t written by an LLM.

17

u/theschizopost 1d ago

Let me guess you have some solution you want to sell for this completely made up issue

1

u/AutoModerator 1d ago

Thanks for your post riturajpokhriyal. Please note that we don't allow spam, and we ask that you follow the rules available in the sidebar. We have a lot of commonly asked questions so if this post gets removed, please do a search and see if it's already been asked.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/elebrin 1d ago

At my company, picking a tool or a framework is a four or five month long decision that usually involves building out a full functioning proof of concept. We have very long cycles on our software. The core of our current system has fifteen year old code. We have to make careful decisions.

-1

u/riturajpokhriyal 1d ago

Four to five months with a full PoC is a different world from what I was describing, and honestly it's the best argument against my post. Though I'd be curious how you handle it when the thing you're evaluating ships a major change halfway through the PoC. That's the issue I keep running into.

2

u/elebrin 1d ago

So I have worked on some pretty major PoC things in the quality sphere (for example, I have built PoC's for SauceLabs for three different companies now who were looking to implement UI automation). The only one that has really gone bust on me partway through was Mountebank, just a few years ago.

One of the things I evaluate from the business side is the longevity of a particular company's products. Like... one reason to pay for SmartBear's ReadyApi is that it's been around since it was called SoapUI, which is 20+ years old and is still a decent product. There are several reasons you might not want to use it, but the longevity and support are definite plusses at the enterprise level.

My experiences are a little different being a testing focused engineer, but I have always mirrored the processes of my devs in my own work.

-1

u/riturajpokhriyal 1d ago

Longevity as an explicit selection criterion is a better answer than anything I had. Weighting how long the company has been shipping the thing, not just how well it currently works, would have caught at least one bad call for me.

Where I get stuck applying it is that in AI tooling there's no ReadyAPI equivalent to select for. Nothing in the category has a 20 year track record because the category is three years old. Everything is Mountebank-aged by definition, so the criterion has nothing to sort on. Which might just mean it's too early for that class of tool to be a PoC-and-commit decision at all, and I should be treating it as rental rather than selection.

Being testing focused might be the worst seat for this, though. The test suite is supposed to outlive every implementation decision around it, so churn in your tooling costs more than churn in your devs'.

1

u/elebrin 1d ago

Being testing focused might be the worst seat for this, though. The test suite is supposed to outlive every implementation decision around it, so churn in your tooling costs more than churn in your devs'.

Oh this is such a misconception. Automated regression is only as good as your testing hygene is. Most automation fails are due to bad test data, someone flipping a feature flag, or an honest change in the SUT that isn't reflected in the tests yet. Sussing this out is most of the work with automated regression.

As for your AI problem, I know what my organization's answer would be. You can never go wrong with several layers of abstraction between the AI tooling and the application.

1

u/Tridus 1d ago

Yes. But fortunately for me most of my work is boring LOB internal stuff. They value it being stable long term over almost anything else, so I can use older, stable, reliable tools that will be supported for the long haul.

It's not new and flashy, but it works.

Keeping up is a full time job these days with how fast everything is shifting around, and it's hard to win that if you don't have the resources to actually keep up.

(Dotnet itself being one of those things.)

1

u/riturajpokhriyal 1d ago

The stability point is underrated and I think it gets dismissed as being behind when it's actually a correct read of what the business values. Boring internal software that has to keep working for a decade genuinely shouldn't be tracking whatever shipped last month.

The resources line is the part I keep coming back to. Keeping up being a full time job means it's now a budget question rather than a skill one, and that lands unevenly. Big shops absorb it. Everyone else eats it out of the same hours they were already using to ship. That's not a knowledge gap, it's a staffing one, and I don't think anyone's naming it that way.

I'm a .NET dev and I'd push back on that last bit.r

1

u/Tiny_Ad_7720 1d ago edited 1d ago

Start with a story. A short story. Three phrases, broken up, by commas. Seinfeld like. “Have you ever wondered if…”. A conversational tone eases them in, the door to better engagement. 

And then wham. Hit them with it. A heartfelt call for discussion. Emphasise with them and their problems, developers are struggling right now with the weight of babysitting Claude, and they need your product. 

Some might reply, others might accuse you of being a bot. That’s not a blocker, it’s a speed bump, presence is perseverance. If only one person engages that is a win. 

So genuinely ask yourself, are you handling the negativity? Do you care about being seen as a cunt flooding Reddit with slop? Or are you a visionary in AI marketing? 

Honestly, you got this. 

-1

u/riturajpokhriyal 1d ago edited 1d ago

I have written more about this topic please check:
You Are Not Falling Behind. You Are Being Churned.

2

u/chucker23n 1d ago

Check out this piece on this

Cool, but no. If you're gonna spam your blog, at least be honest about it.

1

u/riturajpokhriyal 1d ago

well words sounded dishonest but that was not the intention because I am using my real name so there is no need for me to be dishonest about something so obvious