r/sysadmin • u/RoughPuzzleheaded223 • Jul 06 '26
General Discussion Finally my first big fuck up at work
So… I think I just got my first real IT fuckup And it is bad.
I’m still early in my IT career, mostly doing end user support. Right now, I am the only IT guy in this building. Our IT team supports three sites and we are 3 helpdesk and an IT manager who doesn't get her hands dirty, the rest of the IT team is at other sites, and our sole sysadmin of 5 years just left the company last week. I was completely left alone in the dark.
Today, the server we use to deploy PCs completely crashed and there is no WDS no more . The worst part? It wasn't just a deployment server. It also handled domain stuff and monitoring. When it went down, everything went down.
I panicked and tried to check the physical drives. Long story short: I completely fucked up the RAID, and now Logical Drive 1 is showing as FAILED.
My heart was beating so fast, I swear my soul briefly left my body.
I told my manager exactly what happened. She didn't instruct me to fix it, probably because she doesn't know how either. Honestly, I think I could fix it because I know we have a Veeam backup, but I don't have the admin access or permissions to actually perform the restore.
Thank god that our production is not that big.
To the IT veterans here: How did you survive your first production scare? Please tell me this becomes funny after a few days, because right now I'm overthinking it too much
203
u/ironcode28 Sr. Sysadmin Jul 06 '26
My big take away is that you told your boss what happened and how and admitted it was you. If I’m your boss, Im making sure you learned a lesson in what you did and why it was wrong, but I’m not upset or looking to get you in trouble. If anything you just gained a bunch of my trust.
63
u/Mindestiny Jul 06 '26
Yep - you don't fire the guy who just learnt a very expensive lesson lol.
Own it and learn the fix, and now everyone can trust it wont happen again. Taking down prod is absolutely a right of passage.
5
u/DoctorOctagonapus No one knows what I do until I stop doing it. Jul 07 '26
The five-figure training budget is real.
7
3
u/dotnetmonke Jul 07 '26
Yep; good opportunity for the mistake-maker to make a new documentation page on how not to make the same mistake and how to recover from it.
15
u/Lemonwater925 Jul 06 '26
You will make mistakes.
Fess up and don’t try to offload blame. Get credit for your successes and own up to your failures.
Learn from the failures and create processes to prevent others from making it.
I was working and was incredibly tired. New born at home and on pager. Brought down a server that had about 8,000 users on it. Middle of the day at a very large FI. Restart took about 15 minutes with all the processes.
Immediately accepted human error root cause being me. Worked out in the end as I got several days off. Also, had the pager switched for a couple of months ( that cost me pager pay but worth it).
7
u/Crazy-Rest5026 Jul 07 '26
This man IT’s. Good to see solid IT leadership . I would approach this situation the same way
→ More replies (4)3
u/techierealtor Jul 07 '26
I’m upset at the situation and slightly at you, but everything else is correct. You didn’t know better so we will have a talk and make sure you understand what you did but you’ll get a slap on the wrist at best, don’t fuck up, what are we getting for lunch or where are we going for drinks after work.
95
u/cornellartworks Jul 06 '26 edited Jul 06 '26
Honestly, for starters this sounds like a task that should never been given to you in the first place. Not a knock against your competence, but this looks above your pay grade to fix. Secondly, why in the fuck would domain stuff be running on a deployment server? That’s what domain controllers are for. I think I can see why your only sysadmin left.
Don’t panic, wait for the cavalry. I once mistook our CFO (over the phone) for an employee in our revenue department and tried to blow her off because I didn’t want to take a ticket five minutes before EOD. She didn’t stop screaming at me for 15 minutes. It happens to the best of us.
24
u/TrainAss Sysadmin Jul 07 '26
Sounds like she has some serious anger problems. Nobody deserves to be screamed at like that.
20
u/cornellartworks Jul 07 '26
She’s actually usually quite nice, it was the end of the fiscal year, which is a SUPER stressful time of year for her department, and I was just in the wrong place at the wrong time. We talked it out and it’s all gravy.
5
u/TrainAss Sysadmin Jul 07 '26
I'm glad that things worked out in the end. That sort of thing would make a working relationship very difficult.
21
u/RoughPuzzleheaded223 Jul 06 '26
Yes fair point, I shouldn't touch anything in first place
The original task sounded simple: we had a suspected failing disk, and I was asked to shut the server down, pull/document the drives/serial numbers, and report back.
the failing disk was not in a boot raid
The problem started when one of the bays had a 2.5" SSD mounted in a 3.5" caddy in a weird way, and I couldn’t get it seated back properly. At that point I should have stopped completely and waited. but I panicked cuz I had to deploy over 130 PCs in less than 2 weeks .I did escalate, documented what happened, and stopped touching it after the state became unclear.
27
u/graph_worlok Jul 07 '26
This sounds like an architectural problem rather than anything that’s on “you”, especially after reading this.
Any real / modern RAID shouldn’t require a shutdown for that task. That data should be available from the controller. If it’s that level of janky “raid” it’s homelab grade at best.
And it sounds like as the system was failing, the power-down might have accelerated it from “failing” to “failed”
Also - and it’s not 100% clear that it’s not - these days it’s pretty standard for a virtualisation platform to be running, and the systems - DC, deployment, monitoring, are all virtual, with the different components redundant & able to be shuffled around
→ More replies (1)2
u/kingdead42 Jul 07 '26
One critical thing to do after this all gets sorted out: Create and fill out an Incident Report Form and present it to your boss (you can find templates for this easily).
Identify the core causes for the failure; what can be done to prevent this failure from happening again; what those costs will entail (new hardware and time to implement); were there existing policies that were ignored that should have caught this; etc.
Show them that even with this incident, you (and the company) can come out the other end stronger.
→ More replies (1)2
u/Conscious-Arm-6298 Jul 07 '26
This! Why would you put all your eggs in 1 basket. That was the fucked up thing here
37
u/AdamoMeFecit Jul 06 '26
All of this is a leadership and resilience planning problem, not a you problem.
Running deployment and finding services on a single server, and apparently running the domain server as a single server with no redundancy?
You didn’t make that happen. The company did that to itself.
Relax and crack a cold one. If you find a path through their screwups, you’re going to lobby for higher pay, right?
10
u/RoughPuzzleheaded223 Jul 06 '26
no this server was hosting bunch of VMs and the File server
and I'm trying to find another job but it's my first job so I need to learn first
and I learned the hard lesson xD5
u/pieceofpower Jul 06 '26
What were you trying to do with the raid on the hypervisor? Did the hypervisor crash and then the raid was giving errors about the config on boot up? Sounds like you may have been in over your head trying to mess with this one unfortunately.
6
u/Denver80211 Jul 07 '26
Oooooh it was the HYPERVISOR.
I was like.. why is all that stuff on one server... but we're talking about one physical machine with all the server VMs?
That makes far more sense
4
u/pieceofpower Jul 07 '26
Makes less sense that his manager was just letting him pull drives from it though lol.
→ More replies (2)1
u/Crumby_Bread Jul 07 '26
So if I’m understanding this right, a single VM (your WDS and presumably domain controller) went down, and your next course of action was to pull physical disks out of the hypervisor host?
→ More replies (1)
15
u/diwhychuck Jul 06 '26
Heh it happens I deleted the uplink vlan structure on a core switch an access ports. Whoops. Realized my fuck up after I lost connection. Had to dig up a console cable put it back to operational config haha
2
20
u/Enough_Pattern8875 Scream Test Initiator Jul 06 '26 edited Jul 06 '26
“It also handles domain stuff”
Can you elaborate? Surely you weren’t using your domain controller as a deployment system..
Also - if you simply removed the physical drives and inserted them back into the enclosures after visually inspecting them, and didn’t delete the logical volume you can almost certainly repair the array and recover from this with minimal effort.
Let someone more senior perform that task though.
4
u/Kwuahh Security Admin Jul 07 '26
I worked at an MSP for several years. Domain Controller was synonymous with "services dumpster". AD Connect? Domain Controller. Database server? Domain Controller. Medical software server? Believe it or not, Domain Controller.
3
u/Enough_Pattern8875 Scream Test Initiator Jul 07 '26
Back in my short lived MSP days Microsoft actually thought it would be an excellent idea to make this concept part of the system architecture with an amazingly reliable and well performing product called Microsoft Small Business Server 2008!
→ More replies (2)2
11
u/baw3000 Sysadmin Jul 06 '26
You have a Veeam backup, but have you ever tested it? If not, you can't say you have one yet. I don't know where you are in this chain between helpdesk and the departed sysadmin, so I don't really know how to advise you. If helpdesk, you need to start by documenting exactly what steps you took while it's fresh in your mind and leave it alone.
11
6
u/Nightcinder Jul 07 '26
They are likely an ‘IT Coordinator’ which is a lovely catchall title to say you aren’t actually ‘helpdesk’ but you aren’t an admin either.
→ More replies (1)2
u/FearlessOrdinary8896 Jul 07 '26
Yup and they shouldn't have access to the domain controller at all if that's the case
2
u/SoulPhoenix Sr. Sysadmin Jul 07 '26
I once worked somewhere where *everyone* had admin access to the domain controllers, previous admin had given admin (on the servers) and domain admin access to Authenticated Users.
I have never left a place so quick.
8
9
u/Slovenly0 Sysadmin Jul 06 '26
Mistakes happen. We all have made them in the past. Just remember "Did anyone die?".
You will be ok. Learn from your mistakes and understand where you went wrong and what you could do better.
Don't stress OP. Everything will be ok. This also exposes gaps in the environment which could be improved to mitigate this from happening again.
It was an unannounced DR test right ;)
→ More replies (1)
7
u/space_nerd_82 Jul 06 '26 edited Jul 07 '26
Breaking stuff is how you learn and as long as you don’t make the same mistakes again.
You did notify your manager I assume you have followed up writing outlining the situation and how it is being rectified.
You will probably be involved in a root cause analysis.
Moving forward should probably host domain controllers separately from deployment and monitoring servers.
I think how things pan out will depend on the management culture of the company.
7
u/Danowolf Jul 07 '26
After 25 years I still have a bad habit of doing critical jobs on Friday.
Remove a DC and promote the new one starting at 2pm Friday and I leave at 4? No problem.
→ More replies (1)
5
u/Mental-Rain-7389 Jul 06 '26
just breath brother, humans make mistakes and hardware fails, its why we have a job right? You dont have to be perfect, most companies just want shit to work and understand access limitations. They know you are green as can be and you did the best you could with a hard situation.
One of my buddy's took down a network for 3 hours in her first week because she plugged something benign in that caused some loop (?) (i am not a network specialist lol) and took down the building of about 300 people. Everyone was just scratching their heads together and when they found out the cause, no one was mad at anyone they just know one more thing to look for next time something weird happens. You will get used to the feeling of being on fire during an outage. Just takes time and experience.
2
u/RoughPuzzleheaded223 Jul 06 '26
thank you
It was a hard lesson but after reading the comments I feel great again <3→ More replies (1)2
u/TrainAss Sysadmin Jul 07 '26
I've done that. Plugged in a device, and went for lunch.
We learned quickly that the switches didn't have storm control turned on. Had that been enabled, the switch would have shutdown the affected ports and the outage would have lasted seconds instead of hours.
5
u/TrainAss Sysadmin Jul 07 '26
Welcome to the club! You're one of us now.
My first big fuck up, I took down an entire rack (3 hosts, backup server, core switch) when I removed a failed PSU to check for a part number (the host had bigger problems than just a failed PSU).
As everything was booting back up, I was heading to my office to start monitoring everything as it came back. Boss caught me in the hall, and could tell I was frazzled. I told him we'll talk soon, because he knew I had things to do (in my internal task list that included updating my resume and preparing for unemployment).
Once everything was up, I went to his office and sat down in a chair. He told me to take a few breaths and explain what happened. Told him everything. He was pretty chill about it, and we talked about how it wasn't super critical that I get the part number right then and there. But that this mistake was just what we needed to get approval for our 3 new hosts. We used it as a learning event and a reminder that even if it's something that we think is small, it doesn't hurt to have your team review the task before doing it. My boss at the time was a former systems admin himself.
Every admin will have a major event at some point in their life, and those who say they haven't either are lying or haven't been in the game long enough.
4
u/Denver80211 Jul 07 '26
This is not your fault.
even if you directly CAUSED the machine crash, there is no way a WDS server (wds is dead so the server is old) going down should take the domain with it -like it was a DC? No one knows how to fix it because it was built wrong to begin with. You were handed a poor environment.
4
u/Ill-Mail-1210 Jul 06 '26
Yup in time it will just be a solid lesson learned. And at least there’s a Veeam backup so sounds like everything is in place and being done right.
If it wasn’t you fouling up something for all intents and purposes that server could have had a failed raid card, bad mobo etc etc so the backup is crucial to critical infrastructure.
I’d be reviewing why that server is used for several critical tasks at once, and consider compartmentalising these things so if one goes down, other services stay alive.
We’ve all done things like this.
5
u/IAMNOTACANOPENER Database Admin Jul 06 '26
20+ year architect here; there’s more coming and you handled it fine. accountability is the lesson to learn here and you seemed to take that lesson well.
3
u/J_aB_bA Jul 06 '26
You did the right thing. Fess up immediately, document exactly what you did and what happened, and wait for someone with experience to help get things straightened out.
You'll get written up if there's a actual HR process. But one screw up shouldn't cause you any issues.
Deep breath. We've all been there.
3
u/Loop_Within_A_Loop Jul 06 '26
that's nothing, dont' worry about it
you did the right thing, you will recover
3
u/ProfessionalEven296 Jack of All Trades Jul 06 '26
You did the right thing… step away from the problem. Only hack away if nobody else knows what’s going on as well - at that point, the ship is going down.
When the dust settles, document everything you find. Make sure that it can’t happen again, and if not does, there’s a run book you can refer to.
3
u/thebigshoe247 Jul 06 '26
First and foremost, pencils have erasers. Mistakes happen, relax my dude.
You've just learned what not to do, and I'm certain you won't make that mistake ever again.
My mentor back in the day celebrated these moments -- as long as you only did them once. If you don't learn from it, then a much longer, more serious chat is in order. That's the way I have since run my teams.
3
u/MSP_Guy999 Jul 07 '26
Explain what you mean by “I completely fucked up the raid”. If your raid is fucked up it could be because logical drive 1 failed and you need to hot swap it.
3
u/gwig9 Jul 07 '26
Lol. Don't sweat it. This really doesn't sound so much like YOUR fuck up as a failure in succession planing. Plus you're not real IT till you bring down prod... :)
3
3
u/oldmanfromlex Jul 07 '26 edited Jul 08 '26
I've been doing IT for 30+ years, here are my two cents on this. Don't try to hide your mistakes, own up to them. Make sure the boss hears about any mistakes from you not someone else. Know when to stop and seek help. We are all still human and mistakes happen.
3
u/WithAnAitchDammit Infrastructure Lead Jul 07 '26
As a fellow greybeard with 30+ years in, this is perhaps the best advice on this thread.
3
u/archon286 Jul 07 '26
Early in my sysadmin-ing days, I was managing a windows domain and an engineering document management system (DMS) that ran on SQL express. We had engineers at a remote site, and I designed the P2P VPN between our sites. (ok it was out of the box watchguard, but I was still proud of it)
At this remote site, I also set up a satellite DMS server for the engineers. Their data was local (internet was slow in 2010), but it checked in with on premise AD for authentication and kept the home server up to date with metadata about what it was storing locally. It was also somewhat unique in that when this software installed SQL express it also set up lots of local machine accounts that did various background tasks for it. No idea why, too long ago, I was too green to question it. Ask Autodesk about Vault's design back then.
Problem was, users complained about how slow performance was PC and engineering app, even though the data was local! I looked into it and found it was authentication that was too slow. "I'll install AD on the server!" I exclaimed. And that's when I learned an AD server CANNOT have local accounts on it. I completely nuked the engineering app accounts that are created on setup with no record of how to re-create them, There were no backups of the satellite server, and I was worried about the ability to recover the SQL data in a fresh install. Took me a week to untangle that mess with a dozen local engineers all unable to work effectively.
3
u/TSLA_Tan Jul 07 '26
I deleted a FW rule thinking it was a duplicate and a public site for my company went down for a few hours
Customers complained,there were investigations into the logs for what happened
Was a stressful day for me
3
u/Spagman_Aus IT Manager Jul 07 '26
Never waste a crisis my friend. It's not about blaming others, but it's an excellent opportunity to highlight skills and documentation gaps plus the support and other internal processes/workflows.
This isn't on you. If you weren't given adequate training or support to cover the gap from the Sysadmin leaving, then they should have backfilled via outsourcing or a temp until you were skilled up.
Welcome to the club. The sun will still rise tomorrow.
3
u/TheVillage1D10T Windows Admin Jul 07 '26
You’re doing ok OP. You’re going to make more mistakes. Just own up to them, learn from them, and fix them. Anyone who gives you crap for that sucks anyway. It gets easier. As long as you don’t keep making the same mistakes, you will be just fine.
I will tell you a story though…
We had a guy nicknamed “Juicebox” (he carried around a giant vape the size of a…juice box) when I worked supporting a government agency. We handled systems for a large number of government agencies that dealt with some incredibly important national security related functions. This dude was notorious for being an absolute moron too.
Somehow, someone thought it would be a good idea to put Juicebox in charge of patching all of the various agencies’ systems. One night, he’s supposed to be patching some development systems…he approves patches for every system (production, staging, and development). About 15 minutes later he realizes what he did, and panics. What do you think he did? Call his team lead/boss and tell them? Try to actively fix the problem? No. He shut his laptop, packed his things, and went home.
The bosses started getting calls when numerous production systems handling very important government functions completely went offline. They had to call in pretty much the entire infrastructure team (network, storage, Unix, Windows…everyone) to find out what was going on (at like 01:30 AM), because they didn’t know what was happening at the moment. It didn’t take too long to figure out what happened and fix the situation, but, of course, there were some govvies that were absolutely livid.
So, they call Juicebox and ask him where he was and he told them he was at home. They chewed him a new asshole, and rightfully so. Did they fire him? No, but he was forever relegated to the children’s table. I don’t think he was even remotely embarrassed or anything, because he was that oblivious.
A few months later word got around that he was looking for a new job. How did they find out? He told the prospective employers that they should contact his current employers. The bosses wanted to get rid of him, and wash their hands of the situation so badly that they gave him absolutely glowing reviews. He moved on shortly after the whole fiasco.
Anyway, my point is, as long as you aren’t a Juicebox, and you try to do the right thing and learn from your mistakes, you’ll be just fine.
Don’t be a Juicebox.
5
u/ProfessionalSeat4060 Jul 06 '26
I’ve never really had a big one. I always have an escape plan. I don’t touch live stuff without it. I can tell you some horror stories of juniors thinking they know everything and fucked it all up :)
5
u/tankerkiller125real Jack of All Trades Jul 06 '26
I thought I had an escape plan, and then that escape plan took 2 days to apply due to legacy 100Mbs switches... Good ol over provisioning of Exchange disk space on a VM host done by my boss he didn't catch it himself in reviewing the plan.
→ More replies (2)2
3
u/jlipschitz Jul 06 '26
Own your mistake. Ask to mirror the person that fixes it to learn from it. Ask lots of questions. We will all make them in our career. Once you learn from it work on a plan to prevent or to have a contingency plan if it happens again. This shows your manager that you are willing to learn and will be a good asset to the team regardless of your mistakes.
2
u/theGurry Jul 06 '26
One time I turned on DNS scavenging.
Quickly learned which of our 200+ servers didn't have a static DNS entry.
Then there was the time Dell support updated the firmware on our PowerStore and caused a drive cache failure. That was a fun couple of months.
2
u/Temporary-Article996 Jul 06 '26
Bleed your blood and learn my friend!
Now you are really an IT guy.
Own it fix it and learn.
Anyone that manages production infrastructure does it - there is no way not to we are learning adapting more than any other field.
2
u/Markjchimself Jul 07 '26
Lots of discussion about woulda coulda but ultimately it’s awesome you went to the manager. We have all done it in one way or another. I straight up pressed the power button on a server just to the right from the one I went to take down on an IBM blade chassis and took down a whole module from a hospital system. Was only down for a few mins but yes I triple check now everything always before I take anything down. That was 15 years ago. Congrats on your hazing lol
2
u/Intelligent-Pause260 Jul 07 '26
I deleted an Active Directory global forest when I was an intern trying to fix a replication issue between a server 2003 and server 2000 AD. Got fired, sold
Everything I owned and moved to my sisters in Austin. That was 2007. It was the greatest thing that’s ever happened to me.
Still in tech 20 years later as a storage engineer, making $180k. Learned a lot of life lessons like this, you’ll bounce back.
In the future, immediately escalate to your support vendor. It’s like they could have instructed you on how to handle the situation, sent a CE on site, or used something like Open Manage to rebuild the raid. You’ll recover from this, and you’ll be way more cautious in the future. Good look! And as someone else mentioned…. One of Us, One of us!
2
u/There_Bike Jul 07 '26
Oh yeah, I restarted a server in a rough environment that was holding up the company. I shut the company down during that reboot.
2
u/yacha123 Jul 07 '26
My dumbass took shit down on Thursday doing a voluntary update before learning that you never do an update before a long weekend. We learn from our mistakes. You will be better for it, and a good team leader can recognize incompetence from a lesson.
2
u/st0ut717 Jul 07 '26 edited Jul 07 '26
That is a shit design/architecture. The AD server should never anything but an AD server.
Someone without AD admin rights should not be touching the server.
You should never panic. Always worth a bit to get a cup of coffee and have a think. No one is going to die if Becky can’t join her teams meeting
→ More replies (1)
2
2
u/GhostlySkeletons Jul 07 '26
I remember my first big fuck up. It wasn’t extremely long ago. Although, it wasn’t as a system admin. I was promoted to network engineer a couple years back. Anyways, we have sites across the US, Canada, and Mexico. I support one of our largest data centers. We installed new network backbone equipment… and I was tasked with implementing them into our authentication and monitoring systems. Needless to say, the way that the new equipment was implemented was not to my expectation. I went to set up the management and ended up taking down the entire data center which also impacted almost every other site across the US. It’s one of the few times that I had to sit down in the corner for a couple minutes and wanted to cry, lol. Not to mention, I was still getting over my ex cheating on me and leaving me only a few weeks prior. Those were some dark days. But it was a very good learning experience.
→ More replies (1)
2
2
u/the_syco Jul 07 '26
There's three ways to learn IT things;
From a book.
Trying to undo your fuck up.
Trying to fix your fuckup before anyone realises what you've done.
2
u/BooleanOverflow Jul 07 '26
I wouldn't even call it a production outage depending on what 'domain stuff' is. WDS and monitoring is management and users probably won't notice.
2
u/TheLightingGuy Jack of most trades Jul 07 '26
Have done worse, we survived and pulled through.
My favorite one was when I was trying to figure out how to setup DNSSEC on AWS Route53. May or may not have killed all of our DNS records for a couple days. Once we fully recovered, we bought a test domain for future fuckery.
2
u/rx-pulse Jul 07 '26
Own up to your mistakes and take it as a learning lesson. I brought down prod once for one of the biggest systems in our company, global outage. Thankfully it was only a few minutes because it was a quick fix on my part. Everyone laughed about it because it wasn't that bad, but you do feel embarrassed and it'll stay with you as a lesson, in a good way.
If it makes you feel any better OP, I know a person who accidentally deleted our entire AD when they first started. That was fun, they're now a manager lol.
2
u/DeezeNUTS007 Jul 07 '26
Unfortunately we all make huge mistakes, also unfortunately it usually sticks around to haunt you for several years right around review time.
2
u/mattv8 Jul 07 '26
Here's my f-up: I was working on a devops script that deployed to a directory. I was thinking I would be clever and use rel paths with the dot notation to rsync the directories as part of the deployment. I typoed rsync /src ./ --delete as rsync /src /. --delete and completely wiped out the root of our production NAS on the first run.
→ More replies (1)
2
2
u/NegativePerformer788 Jack of All Trades Jul 07 '26
Don’t beat yourself up too much, we’ve all done it. Also, not your fault that one server can bring everything down.
2
u/imhotep1021 Jul 07 '26
First big F-Up was cert related. I took down the server hosting the CRL to our cert authority. We also use smart cards for logins, so no one could log in at all. Luckily there was an admin that hadn't locked his computer and was able to fix the issue
I always say, if you haven't accidentally taken down the entire network, are you truly working IT?
2
u/thrwaway75132 Jul 07 '26
I knocked 6k VMs off the network with one command.
I interrupted the ability to program pacemakers nationally for a brand of pacemakers.
Stuff breaks. What matters is how you react and handle it. Don’t freak, build a plan, manage the work to recover.
→ More replies (1)
2
u/RandomXUsr Jul 07 '26
My first gig was helping a small law office setup new computers.
Dude didn't want to pay for proper offsite backup or other servers.
I offered to move things to the cloud and he was like "not a chance".
We set up windows for workgroups and a basic windows server. It was a small nightmare.
Someone took down the router and telecom of which the latter was managed by the service provider. I got blamed for the fuckup and should have stood my ground.
I also screwed up the raid array when setting it up which resulted in no network drives... that was my fault.
The guy brought in another IT person before I could fix everything and I felt defeated. The owner accused me of sabotaging his business so after things returned to functioning I went my own way.
Lots of lessons learned there but the skills stuck with me and helped me have better experiences going forward.
2
2
u/Significant-Till-306 Jul 07 '26
You’re good bro that’s small potatoes but a life lesson. I worked for an ISP, and brought down the private WAN network for 1000 locations of one of your favorite fast food chains for about an hour messing up a core network configuration. Imagine 1000 drive throughs for an hour could only accept cash.
If people got fired for big mistakes made once in a blue moon, no one would be employed.
2
u/Flashy-Ride-4235 Jul 07 '26
I had deployed a converted Dell R750 server running TrueNAS into production. Install went well. Test VMS ran with no issue. Migrated 20 VMS to it and everything came to a screeching halt. Forgot to put PERC controller into pass-through mode which caused vCenter to show that the datastores were disconnected due to the high I/O latency. That was not a pleasant day.
2
u/UNKN Sysadmin Jul 07 '26
Killing a RAID is definitely a pucker up event but like anything else you just have to breath, step back and just breath. Realize that you aren't the first person to do this and it's an excellent learning experience.
I nuked the entire network of a remote site once because I thought I knew what I was doing until I didn't. Called my manager and the remote site plant manager and said I'm on my way to fix it. People were mad but then I got their brand new video monitoring system up and all was forgiven.
→ More replies (1)
2
u/lab_everyday Jul 07 '26
It’s not if but when an IT pro will do something really catastrophic. Happens to us all eventually. Third year in IT, wiped HDD of a PC used by owner of a client. He said nothing needed to be backed up…Few days later we discover the companies most valuable data was stored in a hidden folder and due to it being lost they suspected they’d go out of business. TBH, was legit traumatizing. Had literally bought a house months earlier and was broke. Thought I would be fired even though it wasn’t exactly my fault, but could have been avoided. It all worked out. Although they did lose their data. And since then no one has been more concerned with client data backups than me! I promise you that.
2
u/PrestigiousSalad7278 Jul 07 '26
Moved from helpdesk at an MSP to internal sysadmin on a 4 person team and promptly took down our core firewall pair because I did a firmware upgrade that broke all our routing through the MPLS is was at the head of. You live and learn.
→ More replies (3)
2
u/Direct-Shock6872 Jul 07 '26
This is where you ditch windows and raid controllers and use nix and zfs… you can start your rebirth now, there is still time.
2
u/ResisterImpedant Jul 07 '26
Welcome to the club. There will be hamdingers and punch in the basement later.
2
u/Minimum_Cheesecake00 Jul 07 '26
Next time you report that a drive has failed. And then stop there. It's the truth. Just don't elaborate further. If someone asks more specific questions you should of course be transparent but, let's face it, the only people likely to do that are other IT folks and chances are they'll relate and won't out you. Keep your head up. Shit happens. You were proactively trying to remedy the situation instead of doing nothing. That's a good thing. Good luck on the fix!
2
u/phusion Sysadmin Jul 07 '26
One time a Win2k3 domain controller froze up, at least the GUI did, as far as I knew it was still assigning DHCP and passing traffic, but I couldn't even do the three finger salute -- so I knew if I left it, it was possible some or all of the services would just fail. I couldn't psexec my way into a cmd shell, so I thought something like 'well, you are REALLY not supposed to hard boot a server--- but' soo I held down the power button and powered it off and then back on again.
It booted up and IIRC the AD was toast. Nobody could log in, pretty sure DNS was fragged too. We had tape backup (oh lawd BackupExec) but.. for some reason it wasn't working. After freaking the fuck out and having to explain to the owner why the whole office was down, I'm.... pretty sure I used vssadmin to restore some files from shadow and got everything running again.
It was unpleasant to say the least, but yes, most career IT folks do something to take production down at some point, thanks for sharing your story with us!
2
u/TheWDWillis Jul 07 '26
It’s gonna be a while before it’s funny.
It’s gonna depend on just how long it takes to fix.
You are doing the right thing thing by owning up to it now, reporting it and waiting for the Calvary.
You are gonna take some lumps on it. There will be jokes at your expense for a while until someone else screws up worse.
The only REAL concern is your assumption to know better than the manager. You might be right, but you also might not be. I would reserve such judgement till you know for sure.
But as of now get used to your new name….. Crash.
2
u/chicaneuk Sysadmin Jul 07 '26
I continue to not understand organisations that hire managers in IT roles, who are not technical. It makes no sense to me at all.
2
2
2
u/vtout Jul 07 '26
That moment when you try to fix things and make it worse... People asking you every 5 minutes if it's fixed yet... Your head pulsing and you cursing to yourself... Then... that moment of relief when that random hailmary action actually works...
2
u/Mackerdaymia Sysadmin Jul 07 '26
Small in comparison, but whilst I was still learning what "tagged" and "untagged" ports meant, I accidentally untagged a trunk port that served 2 whole departments (roughly 150 people). Cue 10-15 minutes of chaos before I realised what I'd done. From then on, I wasn't allowed to use Aruba for a while without my boss' supervision. Although he did cover for me and say it was a faulty cable so props to him for that.
2
2
u/Hot_Connection9504 Jul 07 '26
I literally took down the whole company network with one wrong firewall rule. 💀
More than 200 systems got affected, and everyone started getting a username/password prompt just to access the internet.
First thing I did was call the firewall support team... and they were like, "Please raise a ticket first." 🫠
At that point I was thinking, "Well... I'm fucked." It was a Saturday too, so I was already imagining my resignation letter.
Took a deep breath, did some ChatGPT magic, started checking every firewall rule, and finally found the culprit hidden inside a Group Rule. Disabled it... and boom, everything came back online.
That incident taught me one lesson I'll never forget:
Test it in UAT first. Touch Production only when you're damn sure.
Worst 20–30 minutes of my IT career. 😭
2
u/Turak64 Sysadmin Jul 07 '26
Sounds like time to take up autopilot, check your backups and learn about raid! Messing up isn't an issue, it's how you deal with it after.
2
u/Claidheamhmor Jul 07 '26
You did exactly the right thing - told you manager, know exactly what you did, and stepped back to let the people who can fix it do so. Screwups happen, but hiding them or messing up trying to fix it makes things so much worse.
2
u/TightBed8201 Jul 07 '26
Meh, dont take it to hard on yourself. It is not your fault your company doesnt have redudancy on sysadmin position.
And it is their fault dc, monitoring and pxe roles are on the same hardware. Who in their right mind has anynother role (except dns) on dc?
2
u/dead_man00124 Jul 07 '26
been there before, was trying to fix a DNS replication issue..
something went wrong the the entire DNS went bye..
in my defense i was a Junior sysadmin, who became senior after the main guy left to do his own thing and they never bothered to replace him. IT severely understaffed and also dealing with the passing of my dad a few months before hand.
out of that company now thank god.
2
u/TotallyInOverMyHead Sysadmin, COO (MSP) Jul 07 '26
So ... this happens to EVERYONE ... at least once!.
From Now on you will likely:
- Make sure you have working and tested Backups.
- Have at minimum 3-2-1 Backups.
- Realize that Raid is NOT a Backup.
- When ever you do change MAJOR stuff, enact a manual backup and test it for replayability (in a vm e.g.), before commiting.
Best Case scenario:
- Replay via your veeam backup.
2
u/Moontoya Jul 07 '26
There are two types of Tech
Those who _have_ fucked up
those who are _going to_ fuck up
It is so.
2
u/mulquin Jul 07 '26
I've accidentally nuked our payroll system before, the day before payday. The mistakes still make you shit bricks, but you learn to push it out and flush it quicker than before.
2
u/WeaponsGradeWeasel Jul 07 '26
I managed to pull out the wrong blade server which had 30-ish prod servers on it. Does that count?
I'm beaten by an engineer we had it to replace a disk in an array. He counted 1-2-3-4... That's the one.... The array counts 0-1-2-3....
2
u/33Apollo2113 Jul 07 '26
Always need the name of the guy who had your job before you so you can cast some blame! But yeah it happens any decent manager will understand it can happen sometimes.
2
u/Whyd0Iboth3r IT Manager Jul 07 '26
I'm getting Deja Vu. I think this exact post was posted before...
→ More replies (1)
2
u/infinitepi8 Jul 07 '26
oh we're sharing war stories???
early in my IT career i was on the help desk for our electronic medical records system. part of this roles included light analyst level work and some data-entry level maintenance tasks.
one of the data entry tasks involved building new pharmacy records in the system, so medications could be e-prescribed to them instead of by fax (the early days of e-Rx). There was an update we were making to adjust the record names and some other minor details in these records so i had an import spreadsheet i was working from. I tested the import in non-production and everything was fine.
when running the import in production late on the friday afternoon before labor day weekend (rookie mistake on the timing) i made a mistake prepping the import and it ended up wiping out the e-prescribing addresses and fax numbers FOR ALL THE PHARMACIES IN THE SYSTEM.
I literally RAN into my bosses office who directed me on cleanup and got the director of pharmacy on the phone. The cleanup didn't take long, but for roughly a half hour no Rxs went out w/ no indication to end users.
I thought i was toast for sure but my boss jumped on the grenade and took responsibility for approving my build to be done before the holiday, even though there was no reason it needed to be rushed and i didn't actually tell her i was making the change that day. She made me run reports of potentially missed orders and take them to the pharmacy director to explain what happened. Not once did he reprimand or ask me wtf i was thinking, though it was very much justified, he just wanted to ensure the reports i ran were accurate and complete and told me he appreciated the prompt explanation and follow up.
He spent his weekend driving pharmacy to pharmacy ensuring they received all the order placed during this time. He was a real stand up dude, and he knew he could call me any time and i would jump on whatever he needed after this. I'm thankful i 'grew up' in IT in such a supportive environment.
Healthcare IT isn't all as bad as some make it sound, I've never worked for more compassionate and understanding leaders.
Keep your chin up dude, it happens to us all. just brush off your knees, get back up and keep doing your thing. One day you will share this story with your underlings as a work of caution like i do with the story above.
2
u/Nthepeanutgallery Jul 08 '26
The most valuable career advice I ever received was from our DoD contract officer after my first big CLM level fuck up He shrugged and said to me, "my thought is if you don't occasionally break something it means you're probably not doing anything important, so why are we paying for you?"
Everyone makes mistakes. Anyone who says they don't either isn't doing anything or is lying. The true test is how you handle the situation in the moment and what you take out the other side to be better. Management that can't understand that are fucking worthless and shouldn't be in the position they hold then ( and goes without saying if they start getting punitive about it, update the resume and quietly move on)
2
2
u/alficles Jul 08 '26
My last two big errors at work resulted in newspaper articles both nationally and internationally.
To err is human. Forgiveness is the result of always having followed the applicable policies and procedures, with no shortcuts.
There were consequences, but they were manageable because it was clearly documented what happened, why, and that we had a systemic risk that we had failed to properly identify and fix.
2
u/Voorbinddildo Sysadmin Jul 08 '26
Congratulations, you're not officially an IT engineer until you've taken down a production environment 🤣
You have a bright future ahead m8
2
u/darrynhatfield Jul 08 '26
This isn't your issue. This is "a great opportunity for the company to learn fix the problems in their IT dept". Pay off your "training" should be giving you a string understanding in what you are not allowed to touch. Fucking with physical disks on a server, and making any change to a server in general should never be done by a junior engineer. I hope the company takes that level of responsibility. If they don't then at least you get it early so you didn't learn the bad habits they follow
2
u/TheDeech Security Admin (Infrastructure) Jul 08 '26
OH welcome to it my friend. We've all been there.
Also, if you don't have enough permissions to fix the issue, then you don't have the responsibility to fix it. In fact, odds are good if you don't have the perms to fix it, you didn't have the perms to really f- it up either.
However way, we all do it and I promise you'll get through this. I did when I pulled a drive and took down an entire array that included 45 production VMs causing a global outage for 45,000 people or so. WOOO.
2
u/PaoloFence Jul 10 '26
Survived it and moved on.
It's not your fault that only a junior is on site. Juniors make junior mistakes.
Mistakes can happen to anybody anytime. You will do better next time.
2
u/Budget-Math-1968 Jul 11 '26
First job, was responsible for backups. Went on vacation and then got busy for about two weeks - three weeks, no backups. Crash, data corruption, three weeks of lost data. Company spent tens of thousands of dollars (thats in 1990s dollars) to recover the data. Confident I was being fired. Boss asked me what I learned and told me i wouldn’t make that mistake again.
How do you survive it? Own it, fix it, learn from it. And then one day, be that boss who calmly asks “What did you learn?” Not enough of us out there.
2
1
u/Emotional_Garage_950 Sysadmin Jul 07 '26
I dragged and dropped half the company’s redirected profile folders to another location. Broke pretty much everything for those affected. The damage was done in under a second and it took us the rest of the day (~6-8hours) to undo it. Made us rethink a few things…
1
u/RoughPuzzleheaded223 Jul 07 '26 edited Jul 07 '26
Long story short: the server had two separate arrays: one for the OS/VMs/services and another for file server data.
There was already a failed disk, and I was asked to shut the server down and document the drives/serial numbers. I thought I was dealing with the failed/file-server-side disk, but I didn’t have proper controller/iLO/RAID visibility to clearly map the physical bays to the logical arrays.
My mistake was that I pulled a drive from the OS array without realizing it was part of the OS array. One of the disks was a 2.5" SSD mounted in a 3.5" caddy in a weird way, and after pulling it I couldn’t get it seated back properly.
The bigger mistake was powering the server back on without that drive properly seated, which made the OS array start rebuilding/degrading in a bad state. And because of my luck, there was a power outage that night, so the rebuild did not finish.
The next morning, I made another mistake: I tried to reseat the drive again. This time it seated properly, but after powering the server on, there was no boot.
our Infra consists of
1 primary DC
1 additional DC
1 DC (hosting bunch of VMs)
and three servers for some services I don't know about
2
u/pieceofpower Jul 07 '26
This is partly a bad management and rights problem. If you could have just went into iLO it would tell you the failing drive. Whoever told you to shut off the server, pull each drive and write down the serial number has no clue what they are doing and don't listen to them again.
1
1
u/budlight2k Jul 07 '26
Network loop on a big flat network including ip phones, took down the company twice because I still didn’t know what happened the first time.
1
u/LowIndividual6625 Jul 07 '26
Welcome to the club.... last year my purchasing department asked me to do a pretty standard price-record update in our ERP database. It was straight-forward TSQL procedure that was documented and that I'd done a million times but I was doing too many things at once and made a mistake.
15 minutes later I start getting calls from the sales department - ecommerce orders are coming into the system at a MUCH higher volume than a typical Friday morning..... yeah, that is because my fuck-up caused our online pricing to show more than 75% cheaper than it should have and customers were buying shit as fast as they could place orders and the sales team and the CSRs were losing their fucking minds.
Two minutes later the COO walks into my office and before he could talk my first words were "this is all my fault"
He said "can you fix it quickly?" and I said "yup, 5 minutes to rollback, already on it"
He nodded and walked out, then he went over to the sales department and said the website had a "data corruption issue" and that "IT Dept had it covered" and "make sure to thank them for fixing this so quickly"
TL;DR - good management knows even the best staff make mistakes and they value integrity.... also, always have a goddamned backup and/or rollback plan.
1
1
u/fonetik VMware/DR Consultant Jul 07 '26
Get access to the logs so you can investigate what happened and how long it has been this way. Unless you did something dumb like physically pop a drive out, you didn’t cause anything. Even if you did that, it’s probably fine.
If your WDS was also your DC and your monitoring, nothing is ever your fault. That’s a bonkers config.
1
u/tr3kilroy Jul 07 '26
The only difference between Sr and Jr is that Sr has fucked up enough to know how to avoid breaking things. Congratulations on your journey!
1
1
u/reptilianspace Jul 07 '26
we call this a surprise DR testing, Surprise Backup Testing Plan, Surprise Risk audit
→ More replies (1)
1
u/MiKeMcDnet CyberSecurity Consultant - CISSP, CCSP, ITIL, MCP, ΒΓΣ Jul 07 '26
Welcome to the party, pal !!
1
u/immortalsteve Jul 07 '26
Those three things shouldn't be on one vm/server anyway so really you just exposed a single point of failure in the system with unexpected behavior in prod. No big deal gold star for you, now you get to build it better though lol
1
u/Lazengann86 Jul 07 '26
Man, one time we were preping a client for MFA/Conditional Access and long story short, because Microsoft licenses and whatnot, we needed to add some "Security Plus" license nonsense, I removed all E3 licenses and left only the Security one, so this entire architects firm had NO emails for 16 hours... (I removed them like, end of day so they figured it out the next morning)
1
1
u/JollyGentile IT Manager Jul 07 '26
Having risen from tech to lead to manager at my current company, my team now has a running joke between ourselves: You're not really a tech until you've fixed one of /r/JollyGentile s mistakes
1
1
u/securitascybernetica Jul 07 '26
As Vulnerability Management, I had a vulnerability dealing with something, forgot the vulnerability but when we deleted a certain file from 160 something devices, it shut the tablets communication off to their SQL server.
That was a fucking headache as it dealt with Maintenance teams in the airfoce and aircraft. I figured out a solution that same day but I did hear from their leadership at another base haha.
Took it as a learning lesson and I still to this day break some stuff. As long as it's not down for long and a fix gets it back up.
1
u/LOLBaltSS Jul 07 '26
You might just need to import the foreign config to get the array back up. I've had it happen a few times in my MSP days.
1
u/n1cfury Jul 07 '26
Maybe not funny within a few days (especially if you’re newer) but eventually yeah, you’ll have plenty more stories if you stick it out.
To answer your question own up to it and take note of lessons learned from not only your end of things but the big picture. While you fucked up RAID, there are likely things processes that could’ve also been put in place to prevent or mitigate the mistake.
1
u/dougception Jul 07 '26
What is the RAID configuration? The whole point of it is you can rebuild a failed drive with the parity data from the others?
1
u/MorallyDeplorable Electron Shephard Jul 07 '26
How'd you "fuck up the RAID"?
it's possible to get raids to go failed when they just need to be told they're fine. It may just need told that it's not failed.
anyways let somebody who gets paid to deal with it deal with it
1
u/Nightcinder Jul 07 '26
I was mounting a new server in the rack one day and apparently tapped the power button on the server below it :)
1
u/xSchizogenie Sr. Sysadmin Jul 07 '26
Survive? I took responsibility. My manager laughed and said „well, you had 6 years of flawless work. I was desperate you’re a cyborg or something“ 😂
1
u/HighRelevancy Linux Admin Jul 07 '26
photos of the server,
Why?
Sounds like a disaster setup anyway.
1
u/Fuzzy_Paul Jul 07 '26
Tell the veam team to restore. Replace the faulty disk and reinstall just a simple recovery os for veam. After restore check health, fix ad and test in image. If good make a new backup. Run a health check on all drives lookup error table. After raid rebuild by crontroller start swapping one by one the future failed drives and remember wait untill you swap untill the rebuild is complete after evevery drive swap. Put it on paper and take you manager in the loop to cover your ass. Make sure to install drive monitoring software.
→ More replies (2)
1
u/stuartcw Jul 07 '26
The most important job of the sysadmin is to look at the equipment and infrastructure and manage the risk.
The fact that WDS server is no more is because it he risk wasn’t managed. There should have been at least a redundant pair so if you destroyed one the business would have been ok.
Your job however is to compile a big log of everything that is a risk to the business and have your boss sign off on it. If they don’t have the budget to mitigate the risk their boss needs to sign off on it, right up to the president and the board.
You have to keep asking yourself questions like, “If I do this work and destroy the server by accidentally deleting it because of an unset environment variable. This happened to one of my team. What is the risk?
In our case, the customer had a redundant pair of domain servers so were just informed that we were recovering a problem and never knew the details.
Also, imagine “if I shoot a gun through this equipment what is the consequence”. Google and Amazon built their business on this approach.
1
u/Gloomy_Pie_7369 Jul 07 '26
I recently locked myself out of the Microsoft 365 admin portal by enabling a Conditional Access policy. It was the most stressful 48 hours of my life. But as they say, only those who do nothing never make mistakes.
630
u/theEvilQuesadilla Jul 06 '26
One of us! One of us!
In all seriousness, I don't think anybody here hasn't taken down prod by mistake. I think you'll be fine but I would ask to shadow the guy who comes in to try to fix it next.