Potential serious issue with CTA not coordinating mounts and/or sending two commands to the same drive

I have been trying to gather more information on an issue Fermilab has encountered with CTA on two different libraries, and that Spectra has reported as a known problem at another CTA site (I do not know which site).

In May, we experienced a problem with a tape header being overwritten on an LTO-10 drive in an IBM TS4500 library which uses a SAS switch. Initially, it was thought to be a tape defect. IBM’s analysis concluded the following:

Good Morning Tim,

Update / analysis from our developers:

This appears to be an application / host / usage issue, not a drive or media issue. The tape was \[over\]written from BOP without any label records on a prior mount/drive. The very first logical objects on tape are block size 0x40000 bytes. The READ of 0x50 bytes \[seemingly expecting a label record\] correctly got an ILI check condition with the sense data indicating that an overlength read was performed (only the first 0x50 bytes of the 0x40000 record at LB 0h was returned to the host as indicated by the sense data). The drive is working as expected given what was written to the tape.

Perhaps there was something that happened during the \[most recent\] writing mount, where it appears the application overwrote the tape from BOP starting with data (overwriting the label and likely other data that was on tape \[written on prior mounts\]). The tape has 1039 filemarks and 6658700 records \[presumably 0x40000 bytes each for \~1.75TB\], so it seems like much of the 1900 files / 10.6T were overwritten. Perhaps a rewind or other command to position to BOP was issued during writing on this mount. Any further debug needs to focus around the last writing mount (written by canister S/N 781357D (drive S/N YB1097000683) on 2026-05-20). Such debug would likely need application / system logs and/or a drive dump around this time.

To avoid / detect overwrite issues, it is recommended that applications use best practices including: using reservations, checking position before/after writing, using append only (aka data-safe) write mode. This is especially recommended in multi-initiator environments. There are two or three initiators on this reading drive on a single SAS port (not sure how the writing drive was configured). It seems like there is some kind of SAS switch in this configuration. Perhaps some initiator \[other than the writer\] issued a rewind / load during the writes (reservation or append only mode can protect against this, and position checks before / after writing can avoid or at least detect overwrites at the time of writing).

Unless you have a drive log/dump from drive 781357D/YB1097000683 following the failure on 5/20 I think we may have exhausted our data for further analysis, but let me know if you have any questions.

Thanks,

Ron

After this, I changed the zoning on the SAS switch to only advertise specific drives configured for each host. However, it seems that this may not resolve the problem.

We recently had similar issues on a Spectra library which uses direct-connect Fibre Channel and two drives per host. Spectra provided me with the following analysis:

Hi Tim, My Engineering team took a deeper look at this and found that the problem here is This is seems like a host software issue. Based on the inventory log, we suspect 2 different initiators are asking the same library to move a tape from the same source slot to different drives. CTA does not coordinate the hosts, which is a fundamental flaw in the software. This lack of coordination has led to accidental overwrite of data at a different customer we have using CTA, so it’s a serious problem. You should get with the provider to get help for a solution on this. This could cause serious issues with over writing data. Thanks

Spectra also provided me with log snippets from the library demonstrating the problem.

To summarize: Is this a known issue? Is there a timetable for resolving it? If there is no timetable, is there anything I can do to mitigate?

I think we have tracked this and similar issues to three tapes, and I am happy to provide any logs that you think may provide additional insight. Thank you.

Hi Tim,

Thank you for bringing this issue to our attention with your detailed message.

It does indeed sound like a serious issue, and rest assured, we will treat it as such.

However, before we go any further, we need to fully understand your setup. We would appreciate it if you could provide us with a diagram showing:

  • How the tape servers are connected to the tape drives
  • Where the SAS switch is located in the configuration
  • Where cta-rmcd is running
  • How many cta-taped processes are requesting mounts from each cta-rmcd

If you have different setups for different libraries, please provide individual diagrams for each use case.

If you are short on time, a hand-drawn sketch on a piece of paper would be completely fine.

I would also like to understand the situation with the other Spectra Logic customer using CTA.
Do you have more visibility there?

Once we have your diagram, we can set up a Zoom call to discuss this in detail.
Alternatively, there will be a dedicated CTA worldwide operations meeting on October 1st where we could discuss it.

We are aware of some deficiencies in the CTA tape mounting process and perhaps have not communicated this clearly enough to external sites.

Best regards,

Vladimir Bahyl
CERN