Monday, December 6, 2010

More on “Stratified Sampling”

Last month we covered the example of applying stratified sampling to segments within a given piece of equipment. For this month, I will cover an example where it is used for a series of equipment items in a manufacturing equipment train (for finished drug product manufacture). Let’s assume for simplicity that there are three equipment items used for manufacture of the drug product. I’ll call them P, Q and R. I'm cleaning Product A, and Product B is the next manufactured product. Let’s assume the surface area of the three equipment items are as follows:
Equipment P: 100,000 cm2
Equipment Q: 85,000 cm2
Equipment R: 15,000 cm2
In a typical carryover calculation, I would calculate my limit using the total surface area of the equipment train to arrive at a surface area limit. Let’s say that result (based on the dose of the active in Product A, the dose of the drug product Product B, the total shared surface area, and the batch size of Product B) is a L3 value (limit per surface area) of 1.0 μg/cm2. I would then require in my protocol that each swab sample meet that limit of 1.0 μg/cm2. If I were doing a separate sampling rinse for Equipment Q, and if I used a sampling rinse volume of 50 L, I would set a rinse limit for that equipment item based on conventional rinse calculation:
L4 = (L3) (Surface area sampled) = (1.0 μg/cm2) (85,000 cm2) = 1.7 μg/mL
(Rinse volume) (50,000 mL)
I would then expect my rinse sample to meet that L4 limit. And those determinations are perfectly acceptable (and commonplace) ways of determining that I meet my acceptance criterion.
But, I can also use stratified sampling to determine compliance with my calculated L2 limit (the total carryover). Remember that there are conditions to utilizing stratified sampling in this way. The primary concern is that the residues carried over from equipment surfaces are uniformly distributed in the next manufactured product.
Continuing with the example I started with, the equipment train is stratified by the individual equipment items in that train. Then during my protocol, I want to make sure I sample all the worst-case locations in each equipment item. Note that this sampling is essentially the same sampling (locations and number of samples) as if I were not doing stratified sampling. I then measure residues in all samples. The next step is to multiply the highest value of any swab sample for any swabbed site within a given equipment item by the surface area of that equipment item. This gives me a maximum possible actual carryover for that equipment item. Let’s assume the measurements in the second column below are the highest values obtained for a given equipment item. The maximum actual carryover for each equipment item is given in the last column.
Equipment P: 0.15 μg/cm2 X 100,000 cm2 = 15,000 μg
Equipment Q: 0.20 μg/cm2 X 85,000 cm2 = 17,000 μg
Equipment R: 1.20 μg/cm2 X 15,000 cm2 = 18,000 μg
The total possible carryover under this scenario would be the sum of all equipment items, or 50,000 μg (50 mg, for those of you used to seeing smaller numbers). Going back to my original calculation of a total carryover limit, if my average L3 was 1.0 μg/cm2 and if the total surface area were 200,000 cm2, my total L2 limit would be 200,000 μg. Since my actual maximum residue value as determined by stratified sampling was 50,000 μg, my cleaning was effective because it was under the residue limit.
The value of this approach is seen in that with the example used, I would have failed the protocol (at least for Equipment R) with this data, with at least one location in Equipment R being above the calculated L3 limit of 1.0 μg/cm2. But using a stratified sampling approach, I still have a scientific and logical rationale for saying the carryover is less than the calculated total amount. Yes I would be happier to see all data points meeting the one L3 limit of 1.0 μg/cm2. But, I would also be happier if al
l my data points were below the limit of detection (LOD) by the best available analytical technique. But at least for most situations, that is not required.
Some of you may question the use of this technique, never having seen it before. Frankly, when I started as an independent consultant, I had not seen this “stratified sampling” approach. When I saw it, my response was something like “It’s not the typical method most pharmaceutical companies use, but it does have a scientific and logical basis for use.” Particularly with all the talk about wanting to be on a sounder scientific rationale, there should be no serious objection to this technique provided it is used correctly and in appropriate situations.
Realize that this technique offers an advantage to a manufacturer mainly when a smaller equipment item has a larger residue value. However, it also offers some advantages when the larger equipment item has the higher swab residue values. Reversing the second column values for P and Q in the example previously given results in the following carryover values:
Equipment P: 1.20 μg/cm2 X 100,000 cm2 = 120,000 μg
Equipment Q: 0.20 μg/cm2 X 85,000 cm2 = 17,000 μg
Equipment R: 0.15 μg/cm2 X 15,000 cm2 = 2,250 μg
In this case, the total carry is 139,250 μg, which is still below the acceptance limit of 200,000 μg.
However, it is possible to carry this approach only so far. Here is a third case:
Equipment P: 0.15 μg/cm2 X 100,000 cm2 = 15,000 μg
Equipment Q: 0.20 μg/cm2 X 85,000 cm2 = 17,000 μg
Equipment R: 5.20 μg/cm2 X 15,000 cm2 = 78,000 μg
I can total the carryover to get a value of 110,000 μg, and it looks like I will pass my total carryover acceptance limit of 200,000 μg. However, in this case it is likely that I will fail my visually clean criterion (with a swab value of 5.20 μg/cm2). In other words, the use of this technique is not completely elastic.
The purpose of the Cleaning Memo is to present additional examples of the use of a stratified sampling approach for determining compliance with residue acceptance criterion in a protocol. This approach, while not commonly used, is based on good science and good logic. This Cleaning Memo should be read in conjunction with last month’s Cleaning Memo.

Final Notes on “Stratified Sampling”

The first question should be an obvious one. If stratified sampling can be applied to segments within a given piece of equipment and it can also be applied to a series of equipment items in a manufacturing equipment train, is it possible to apply stratified sampling where I stratify both segments within one piece of equipment and a series of equipment items in a manufacturing train? And the answer is “Yes”, although it significantly complicates the pre-protocol work that must be done, as well as the calculations necessary in the protocol execution itself. However, there is nothing logically or scientifically invalid about applying stratified sampling principles in these cases (provided of course, the restrictions or limitations discussed in the two previous Cleaning Memos are considered).
The second question involves the use of rinse sampling and swab sampling for a stratified approach (the examples given in the previous months all involved swab sampling). I will consider three cases. Case I: Is it possible to use this approach where only rinse sampling is performed (on separate equipment items in a train)? Case II: Or, to use it in an equipment train where some of the equipment is sampled by swabbing and some equipment is sampled by rinsing? Case III: Or, to use it in an equipment train where a given equipment item is sampled by both rinsing and swabbing? The answer to all questions is “Yes”. The key is just to use the stratified sampling principles appropriately.
In Case I (everything is only rinsed), it is important that the rinse limit be established on carryover calculation principles. Then, the total actual carryover for a given equipment item can be determined by multiplying the concentration of the residue in the rinse solution by the volume of the rinse solution. The total actual carryover for the equipment train can be determined by adding up the actual carryovers for each equipment item. That total is then compared to the total carryover limit determined by the MAC calculation.
In Case II (some items in a train are only swabbed and some are only rinsed), it is important that both the rinse limit and the swab limit be established on carryover calculation principles. For equipment items rinsed, the total actual carryover for a given equipment item can be determined by multiplying the concentration of the residue in the rinse solution by the volume of the rinse solution. For equipment items swabbed, the total actual carryover for a given equipment item can be determined by using the calculations shown in the March Cleaning Memo. The total actual carryover for the equipment train can be determined by adding up the actual carryovers for each equipment item (whether it sampled by swabbing or rinsing). That total is then compared to the total carryover limit determined by the MAC calculation.
In Case III (some items in a train are sampled by both swabbing and rinsing), it is again important that both the rinse limit and the swab limit be established on carryover calculation principles. For each equipment item both swabbed and rinsed, the total actual carryover for that equipment item can be first determined by using the swabbing data to determine the maximum actual carryover for that item based on swabbing. Then for the same equipment item, the maximum actual carryover is calculated based on the rinse data (using the principles in Case I above). Assuming both the swab data and the rinse data each give results representing the total actual carryover, the larger of the two actual carryover results is then used for the total carryover for that equipment item. [Note that other things being equal, the results based on swabbing will ordinarily give a higher result than the data based on rinsing.] This is done for each equipment item, and the total actual carryover for the equipment train can be determined by adding up the actual carryovers for each equipment item (however that item is sampled). That total is then compared to the total carryover limit determined by the MAC calculation.
This sounds like a lot of work, and it does represent extra calculations. However, there is another alternative in the use of stratified sampling. That alternative is to use a staged approach to determine whether the acceptance criterion in a protocol is met. This staged approach involves a initial evaluation of every swab and rinse sample result, and comparing it to the acceptance limits for swabbing (perhaps based on a limit expressed as μg/cm2) and for rinse samples (perhaps based on a limit expressed as ppm or μg/ mL). This is how it is ordinarily in a cleaning validation protocol. If all those results are at or below the acceptance limit, then there is no point in going further in stratified sampling; the residue limit criterion is met. However, if one data point (or more) for a swab or rinse sample is above the limit, then the next step is to proceed with a stratified sampling approach to see whether the total carryover is acceptable. In this case, it is preferable not to say the initial evaluation “failed” the acceptance criterion. It is better to say something like “Stage 1 criteria were not met, and we will proceed to a Stage 2 evaluation to determine acceptability”. If the stratified sample approach demonstrates that the total carryover was acceptable, then the residue limit criterion was met.
In this staged approach, one cannot have “unacceptable” results from a typical (Stage 1) evaluation, and then decide that it might pass by a stratified sampling evaluation. This staged approach should be written into the protocol. Furthermore, if segments within an equipment item are to be stratified, those segments should be identified in advance (in part to prevent the temptation to “adjust” the segments or segment surface areas based on the data obtained, so that the end result is more likely to meet the limit based on stratified principles).
This staged approach should not be foreign to those in pharmaceutical manufacturing. It is something we use on a regular basis in the USP testing for conductivity in Purified Water and WFI systems.
The purpose of the Cleaning Memo is to address additional issues in the use of a stratified sampling approach for determining compliance with residue acceptance criterion in a protocol. This approach, while not commonly used, is based on good science and good logic. I should reiterate some conditions for utilizing stratified sampling. This method can only be used if the residues from equipment surfaces are uniformly dispersed through the next manufactured product. Furthermore, it is preferably only used proactively. That is, define the segments in advance, select the worst case location(s) in each segment, and write your protocol with this approach. It is also preferable that this approach be permitted in your cleaning validation master plan or high level policy. Furthermore, this Cleaning Memo should be read in conjunction with the March and April (2010) Cleaning Memos.

Acceptable Variability for Sampling Recovery Studies

Several months ago (January 2010), my favorite statistician, Lynn Torbeck*, published an article in Pharmaceutical Technology entitled “%RSD: Friend or Foe”. The article was basically about the misuse of statistics. In it Mr. Torbeck made the statement that applying a percent relative standard deviation (%RSD) criterion to percentage recovery values in recovery studies was not statistically valid because the values themselves were already percentages. Now, it is common practice in cleaning validation programs to include a criterion for %RSD for recovery percentages in sampling recovery studies (such as swab recovery studies). For example, companies might specify for swab recovery studies that the data collected for a given spiked level have a %RSD of ≤15%.
This got me thinking. Is Lynn right, or are most of the pharmaceutical companies (as well as yours truly, who has taught the use of a %RSD criterion for recovery studies) right? Furthermore, if we abandon the %RSD criterion, what criterion do we use to measure variability in a sampling recovery study? And, finally, does it really make that much difference? In situations like this, I often try to look at practical data to see the impact.
Here is my first example, which in comparison to a second case, illustrates that %RSD is not necessarily a good measure of variability in a sampling recovery study. For simplicity, I am only going to illustrate this with swab sampling, involving one operator who performs three replicates. In Case A below, the 100% recovery value is 2.03 µg/cm2. The values obtained, the average, the standard deviation and the %RSD is given in the Table A below. For simplicity, the data on recovery values will omit the units (µg/cm2).
Table A Data
  Data values % Recovery Values
Replicate 1 value 1.93 95
Replicate 2 value 1.96 97
Replicate 3 value 2.01 99
Average 1.97 97
Standard deviation 0.040 2.0
%RSD 2.1 2.1
Now, we have a second operator with the following data for the exact same sampling situation. Table B has the data for that operator in this situation.
Table B Data
  Data values % Recovery Values
Replicate 1 value 1.03 51
Replicate 2 value 1.06 52
Replicate 3 value  1.11 55
Average 1.07 53
Standard deviation 0.040 2.0
%RSD 3.8 3.8
Now admittedly this is an extreme case. With one operator getting recoveries of 97% and a second operator getting recoveries of 53%, I would suspect something is wrong. However, that doesn’t change the statistical evaluation. If I use %RSD for each operator, then it appears that the variability of the data with operator B (%RSD = 3.8) is much greater than the variability of operator A (%RSD = 2.0). However, when one looks at the data itself, and specifically at the standard deviations, the standard deviation in each case is the same, suggesting the variability in each case is the same.
If it doesn’t make sense to use %RSD as a measure of variability, what can we use instead? If we look at the data in Table A and Table B, perhaps we fall back to using just the standard deviation of the values themselves as a measure of variability. However, while that works in those two specific situations, what will happen in a significantly different case? Let’s suppose the 100% recovery value is not 2.03 µg/cm2, but is 20.3 µg/cm2. Table C has one possible data set for that situation.
Table C Data
  Data values % Recovery Values
Replicate 1 value 19.3 95
Replicate 2 value 19.6 97
Replicate 3 value 20.1 99
Average 19.7 97
Standard deviation 0.40 2.0
%RSD 2.1 2.1
Note that the 100% value in Table C is 10 times higher than the 100% value in Table A. In addition, the data for the three replicates are 10 times higher than the data in Table A. In this case, we look at the actual standard deviation values, we see that they are significantly different (0.040 for Table A vs. 0.40 for Table C), and we conclude that perhaps in this situation the %RSD is a better indication of variability of the data.
What we are faced with is a dilemma. What gives a better indication of variability of replicates, the standard deviation values (which appear a better measure in comparing A and B), or the %RSD values (which appear to give a better measure in comparing A and C).
One possible solution is to base the measure of variability on the standard deviation of the data itself as a percentage of the 100% recovery value. What this does is normalize the data, so that the variation in Table A and Table B are the same, but also the variation in the data in Table A and Table C are the same. The data expressing the standard deviation as a percentage of the 100% value is given below for each of the three cases covered above:
Table A Case: 100(0.040/2.03) = 2.0%
Table B Case: 100(0.040/2.03) = 2.0%
Table C Case: 100(0.40/20.3) = 2.0%
In looking at the data in the three situations, the variability of each operator seems to be the same. This measure (the standard deviation of the values as percentage of the 100% recovery value) appears to reflect that similarity. Furthermore, in cases where the data values might be more divergent, it would also appropriately reflect that greater variability.
Where does that leave us? Am I expecting people to start using this measure of variability in place of the %RSD for sampling recovery studies? Probably not. Part of the reason is that a variability criterion for sampling recovery studies is typically set at a relatively high level (15%-20% RSD), reflecting the high variability of recovery studies. Furthermore, if percent recoveries are relatively high (>80%), the difference between the proposed new measure and %RSD is somewhat minor. Finally, if percent recoveries are low but still acceptable (such as 50-65%), then the %RSD measurement will give a higher measure of variability, thus reflecting a worst case. So, while this proposed measure may provide a more scientific basis for the degree of variability, the existing method is not terribly wrong (particularly for something as variable as a swab recovery study).

Statistics for Visual Limits

This Cleaning Memo is an evaluation of a recent Pharmaceutical Technology article entitled “Statistically Justifiable Visible Residue Limits”, by M. Ovais (March 2010 issue, pages 58-71). The author asserts that “current methods for establishing visible residue limits are not statistically justifiable”. The author presents an example of determining a “visual residue limit” by spiking studies. In a spiking study, coupons are spiked at different levels, and a panel of observers looks at each panel under defined viewing conditions to determine the nature of the residue. Without going into the detail, the author provides a “logistic-regression” model to determine the “probability of detection” of the residue at that selected level. Needless to say, what this results in is a higher limit (a worst-case) than what would be determined by a consensus of multiple observers.
There are several questions to ask. Is this statistically correct? And, is this statistical evaluation really necessary? I can’t answer the first question; I’ll leave that up to the statisticians. While one answer to the second question is “you can certainly do it because it results in a higher visual limit” (a higher visual limit, contrary to what is often thought, is actually a worst-case), I’ll give my answer to the second question below.
However, to do that it is necessary to clarify a few things. One is that many publications (apparently including this Pharm Tech article) list the “visual limit” as the lowest spiked level at which observers can consistently see any residue on the spiked surface (of course this is under defined viewing conditions and for a defined residue and a defined surface, but that will be assumed throughout this discussion). In other words, if I spike a surface of 25 cm2 at different levels, then the spiked level at which all observers see even a speck on the spiked surface is the visual limit. While that may be one definition of visual limit, it is not a useful definition for cleaning validation purposes.
Why do I say it is not useful? The main reason is that the purpose of a visual limit is to say any surface viewed (under the same viewing conditions) that is visually clean has residue below that defined visual limit. Unfortunately, doing spiking studies and determining the lowest spiked level which has any residue on the spiked surface can’t be used in that way. Why? First, remember that the worst case for a visual limit is a high value, not a low value. Defining the visual limit in this way presents an artificially low visual limit, which will allow one to state that the residue is below the specified value without a sound scientific basis.
The issue here is that when I spike at a fixed level (let’s say 1.0 μg/cm2), and only see a small speck of residue in a corner of the area spiked, I cannot really say that any surface that is visually clean has a residue level of less than 1.0 μg/cm2. If I spiked at 1.0 μg/cm2, and the surface was evenly covered (an ideal situation), then it would appropriate to say that the spiked surface truly represents 1.0 μg/cm2, and therefore any surface which is visually clean has a residue level below 1.0 μg/cm2.
What happens in the real world when I do spiking studies is that the residue is not evenly spread over the spiked surface. Instead, due to the drying effects (difference in drying between the edges of the spiked residue solution and the center of the spiked residue solution), I will typically see a “donut hole” effect, with differing amounts of residues on different parts of the spiked area. Therefore, if I spike at 1.0 μg/cm2, it is possible that some portions of the spiked area may have concentrations of 0.8 μg/cm2, while other portions have residue levels of 1.2 μg/cm2. Perhaps I can see the residue in the spiked areas where the surface concentration is 1.2 μg/cm2, but not see it at a surface concentration of 0.8 or 1.0 μg/cm2. In that case, I will say the visual limit is 1.0 μg/cm2, which would be misleading.
The “correct” way (or at least one correct way) to determine the visual limit is slightly different. The same spiking coupons are prepared. However, the visual limit is then defined as the lowest concentration (in μg/cm2) in which the entire spiked area is visually dirty or soiled (that is, the lowest level at which residue is seen across the entire spiked area). Defined in this way, there is then a scientific or rational justification for saying any surface observed that is visually clean has a residue below the spiked level. Note that in this case, there may be (or better stated, there will be) variations in amounts of residue on different parts of the spiked coupon. That is inevitable, because of the drying effect. However, this approach is one that should be used (and not the approach of defining the visual limit based on the lowest spiked level where any residue is seen).
Why am I explaining this to address a statistical question? First, there is a certain level of “safety” already built into the determination of the visual limit. When I spike at 1.0 μg/cm2, and state that the visual limit is 1.0 μg/cm2, the true visual limit is lower than that (due to the drying effects mentioned above). How much lower, I can’t say for certain; however, I suspect that the “true” visual limit is probably 0.8 μg/cm2 or below.
Secondly, defining a visual limit is not necessarily an exercise where I need to get that visual limit as low as possible. My preference in using “visually clean alone” is not to do a series of coupons spiked at different levels. I prefer to first calculate the residue limit (using traditional maximum allowable carryover calculations, for example) to determine the limit per surface area (for those of you who follow my writings, this is the L3 limit in μg/cm2). If my calculated residue limit is 4.0 μg/cm2, why do I need to establish a visual limit that may be as low as 1.0 μg/cm2? In this situation, I prefer first do a spiking study at 4.0 μg/cm2. If at that spiked level all observers were not able to see residue across the entire spiked area, then who cares what the visual limit is? I clearly cannot use visually clean alone in a protocol to establish that the residue is below the calculated limit. On the other hand, if I spike at 4.0 μg/cm2 and can readily see residue across the entire spiked area (albeit uneven amounts on different portions of the spiked area), then I have a rationale for saying that surfaces observed that are visually clean are, in fact, below the calculated limit.
Note that this last situation (of spiking at 4.0 μg/cm2) already has some (undefined) safety margin in that the “true” visual limit is somewhat lower (because of the uneven concentrations across the spiked surface). That said, my preference is to add an extra margin of safety. If the calculated residue limit is 4.0 μg/cm2, my preference is to spike at an additional lower level, such as 3.0 or 3.5 μg/cm2. If I can see residue across the entire spiked area at those lower levels, I have an additional margin of safety in my determination of a visual limit.
If what I have discussed is the proper way to implement determinations of visual limits for a use of ‘visually clean alone” (that is, without swab or rinse sampling), then it would appear that there are significant safety margins built into the evaluation, and that a statistical evaluation of the “probability of detection” may be nice to have, but is not necessary.
This leads me to my last point. In that same issue of Pharm Tech was a short article by Lynn Torbeck (my favorite statistician, as revealed in last month’s Cleaning Memo) entitled “The Role of Statistical Tests”. In it, Mr. Torbeck points out that statistical significance tests should only be used after one first determines that there is a practical difference between two data sets. If there is no practical difference, don’t perform the statistical tests.
While the published article on statistics for visual limits is not strictly on statistical significance, it does invoke statistical principles to determine whether future observers would get the same result (or better said, to set a limit such that there is a higher probability that future observers would also report the same visual limit). I would put forth that with the determination of visual limits properly done (that is, defining the visual limit as the lowest spiked level where all observers see residue across the entire spiked area) has sufficient safety margins (either inherent in the process or which can be added to the process) such that extensive statistical analysis adds little or no value.
Note that it is certainly possible to use the statistical approach to further make the visual limit higher (which is a worst case). However, I think a good understanding of what is involved in determining visual limits suggests that there are practical safeguards already built into the visual limit determination.

Visually Clean and Visual Limits

First, let’s clarify the purpose of a visual limit (VL) determination. VL is typically expressed in units of mass per surface area, such as µg/cm2.The purpose is to define a level at which a defined residue is clearly visible on a defined surface (typically a certain material of construction and surface roughness) under defined viewing conditions (typically distance, lighting and angle). The VL is then used to in this way: If the same surface is viewed under identical or more stringent viewing conditions and is visually clean, then the residue is present at a level below the VL. If the VL is equal to or below the limit established by a carryover calculation in a cleaning validation protocol, then the viewed surface meets the defined acceptance criterion without the need to perform swab sampling.
The relevant section in PIC/S PI 006-03 is section 7.11.3, which states:
Carry-over of product residues should meet defined criteria, for example the most stringent of the following three criteria:
(a) No more than 0.1% of the normal therapeutic dose of any product will appear in the maximum daily dose of the following product,
(b) No more than 10 ppm of any product will appear in another product,
(c) No quantity of residue should be visible on the equipment after cleaning procedures are performed. Spiking studies should determine the concentration at which most active ingredients are visible[emphasis added]
The question that we will address is whether for any use of a visually clean criterion, must I do spiking studies to determine the VL? In other words, most people agree that if I am using visually clean alone without any swab or rinse sampling for that surface, then I should do spiking studies to determine what the VL is. While it is commonly stated that visual limits are on the order of 1-4 µg/cm2, we all realize this variable. And further, if the calculated carryover limit is 0.1 µg/cm2, it is not likely that I will be able to use visually clean alone. Furthermore, if the calculated limit is above a certain value (which will depend on the viewing conditions), I would make the case that spiking studies are not required. For example, if the carryover limit were 13.5 µg/cm2 and a stainless steel surface could be viewed in a short distance (such as 2 feet) under reasonable lighting conditions, I would make case that the VL would be significantly below the calculated limit and spiking studies were not needed. Obviously, some kind of reasonableness needs to be used for this latter case; it might not apply if it were a white residue on a PTFE surface.
But the key question for this Cleaning Memo is whether I should (or am required to) determine the visual limit (by performing spiking studies) if I am only using visual examination to supplement swab and/or rinse sampling for the same examined surface.
There are at least two possible interpretations of the “recommendation” in PI 006-03 in Section 7.11.3.c that “Spiking studies should determine the concentration at which most active ingredients are visible”. One interpretation is that this requirement must be read in context, and that context comes from the phrase “the most stringent of the following three criteria”. That is, the requirement for spiking studies only applies if visually clean is the most stringent of the three criteria. If this is the case (and this is my interpretation), then current industry practices are generally consistent with this interpretation.
A second interpretation is that I must perform spiking studies (to determine the VL) for all cases where I use visually clean as a criterion in a cleaning validation protocol. Does this interpretation make scientific sense? Is there a scientific rationale why this should be done in all cases, and in specific where visual examination is done to supplement swab and/or rinse sampling? Let’s see what value it adds in the latter situation. To consider that, we’ll take a look at several examples.
Let’s suppose for the first example that we are in a situation where the calculated limit for a residue is below the visual limit. For example, suppose the calculated carryover limit is 1 µg/cm2 and the VL (if I were to do spiking studies) is 3 µg/cm2. In my cleaning validation protocol, I measure residues for a given surface by swabbing and get results below 1 µg/cm2. In addition, equipment is visually clean. I pass both those acceptance criteria. But, would a spiking study to actually determine the value of VL make any difference in whether I pass or fail. Yes, it might be “nice to know” that the VL for the residue is actually 3 µg/cm2, but is it necessary? My answer is “No”.
For this same example (the calculated limit for a residue is below the visual limit), let’s suppose that when I measure residues by swabbing the results are below 1 µg/cm2. But, my visual examination shows that the equipment is not visually clean. With those results, I would fail my protocol. Obviously it should be clear that the visual failure is not caused by the target residue (for example, the active), because I have analytical data showing that the active is at an acceptable level. The visual failure is most likely caused by other residues (such excipients and/or cleaning agents). Again, knowing a specific VL for the active offers no additional information or benefit.
There are two other situations in this same example (the calculated limit for a residue is below the visual limit). One is where the swab analytical data is above the acceptance limit and the equipment is visually clean. Another is where the swab analytical data is above the acceptance limit and the equipment is not visually clean. I won’t go into detail, for these two cases, but it should be clear in these situation that spiking studies to determine a value for the VL offers no value.
These last three paragraphs have deal with the example where the calculated limit for a residue is below the visual limit. Now we’ll consider the reverse situation, where the calculated limit for a residue is above the visual limit. For example, suppose the calculated carryover limit is 5 µg/cm2 and the VL (if I were to do spiking studies) is 3 µg/cm2. In my cleaning validation protocol, I measure residues for a given surface by swabbing and get results below 5 µg/cm2, and the equipment is visually clean. Does the fact that I have an experimentally determined VL add anything? Well, you might say that if I did a spiking study, I would know that the amount of the active was below 3 µg/cm2. But what is the value of knowing that? If for some reason, I measured the residue by swabbing and the result was 4 µg/cm2 (thus meeting the analytical limit), then if the VL was 3 µg/cm2, I would fail the visually clean criterion even though I did not perform a spiking study to determine the VL. In this situation, it is not possible to have a situation in which I failed the analytical swabbing limit but passed the visual limit.
To sum it up, in these examples (where I am both measuring the residue by swabbing to compare it to the calculated carryover limit and determining the equipment is visually clean), it is not the case that determining the VL by spiking studies adds any significant benefit to confirming that the sampled surfaces are acceptably clean.